NEWS

Cisco MINT Partner! Learn more →

Automation Services
2026-09-18
11 min read

When the Endpoint ID Wasn’t Enough

A customer’s device lookup returned a real endpoint. It was also the wrong host, so we rebuilt the response around identity before isolation.

Cisco Secure Endpoint
Cisco XDR
Endpoint Security
SOAR

A customer’s security operations team asked us to connect Cisco XDR findings with Cisco Secure Endpoint response. Their analysts could already isolate a device manually, but doing so required them to compare identifiers and context across several consoles. The customer wanted the workflow to carry that evidence forward and automate containment when the target was clear.

The first implementation started with the device GUID in the finding. It resolved successfully and returned a healthy-looking endpoint object. The hostname, however, did not match the asset in the incident, and the device’s last-seen time predated the detection by enough to undermine confidence in the association.

Nothing had failed at the API layer. The identifier was valid, the lookup returned data, and the connector could have submitted an isolation request. It still had the wrong host. We stopped the action and rebuilt the design with the customer’s SOC and endpoint team around a stricter principle: resolving an endpoint record is enrichment; proving that it is the incident asset is authorization for response.

Secure Endpoint identity gate: resolve the device GUID, confirm hostname and user context, then isolate or hold.

Turning a mismatch into a design requirement

We began by documenting what the original lookup did and did not prove. The GUID identified the record returned by the endpoint platform. The hostname allowed comparison with the finding. Last-seen time indicated how current the record was. Agent state suggested whether a containment command might reach the device. User and ownership data added context. No single field established identity or eligibility by itself.

That distinction mattered because plausible records are dangerous in automation. A stale asset association, duplicated device record, reimaged host, or identifier carried from an earlier enrichment step can all produce valid objects. A workflow that equates HTTP success with identity confidence can turn an enrichment defect into containment of the wrong machine.

Together, we defined an identity gate before classification or action. It resolved the supplied identifier, compared the incident hostname with the provider hostname using agreed normalization, evaluated record freshness, and retained supporting user and asset context when available. A mismatch or incomplete identity stopped the response. The workflow would not substitute the incident hostname into the returned object or run fuzzy searches until something appeared to match.

We represented that decision with a small, reviewable payload:

{
  "incident_host": "<host-from-finding>",
  "resolved_host": "<host-from-provider>",
  "endpoint_id": "<endpoint-id>",
  "identity": "match | mismatch | incomplete",
  "action": "isolate | approval | hold"
}

The execution path was equally direct:

incident asset
  -> resolve endpoint identifier
  -> compare hostname and context
  -> check age and agent health
  -> classify asset
  -> isolate, request approval, or hold

A hold was an expected engineering outcome, not a failed workflow. It meant the available evidence did not support an endpoint action. The work note had to explain what disagreed and identify who could resolve it.

Separating identity, freshness, and reachability

During review, we found that several useful fields had been collapsed into one idea of “device confidence.” We separated them because each answered a different operational question.

Identity asked whether the provider record represented the asset named in the finding. Freshness asked whether the enrichment was recent enough for the customer’s response policy. Reachability asked whether the endpoint agent and management connection could receive a command. A device could pass identity and fail reachability. It could also be reachable but fail identity, which was the more dangerous case because the API was capable of acting on the wrong target.

We normalized hostnames conservatively according to the customer’s naming conventions. Case and an agreed domain suffix could be handled without changing identity, but truncated, aliased, or materially different names did not become matches automatically. Where the finding lacked a hostname or the provider omitted one, the result was incomplete, not a forced mismatch and not a pass.

Last-seen time became evidence rather than a hidden filter. The endpoint team helped establish how freshness affected each response branch, and the ticket displayed the observed timestamp and resulting status. We did not publish those environment-specific thresholds in the story or hard-code customer names into reusable components. The policy could evolve without changing the connector logic.

User context supported the comparison but never served as the sole identity key. People share devices, administrators sign in to many systems, and service accounts run on servers. A missing owner did not prove safety, while a familiar owner did not prove that the endpoint was the incident asset. Missing user mapping was reported as incomplete enrichment and evaluated under policy rather than silently guessed.

Agent health and management state remained separate from those conclusions. If identity matched but the agent was offline, the workflow returned an unreachable or pending response state and handed the case back with the evidence intact. It did not label the endpoint a mismatch, and it did not claim isolation merely because a request had been accepted.

Applying customer policy only after identity

Once the identity gate passed, the workflow classified the asset for response eligibility. The customer did not want every endpoint treated like a standard employee workstation. Shared jump hosts, infrastructure servers, executive devices, and systems in active maintenance required different decisions because isolation could affect people and services beyond the incident.

The endpoint and security teams supplied those exception cases. We translated them into explicit policy branches with an outcome, reason, and owner. An eligible workstation could continue to automatic isolation. A shared or high-impact asset could request approval. A maintenance exception could hold or route the action according to the team’s current policy. An unknown role did not default to “workstation.”

Keeping classification after identity prevented policy from legitimizing the wrong target. There was no value in deciding whether a returned device was a protected server if we could not first prove that it was the device in the finding. When identity failed, the workflow stopped without evaluating containment eligibility as though the target were established.

We also kept policy data outside the connector-specific lookup. That let the customer adjust which roles needed approval without rewriting the Secure Endpoint interaction. It reduced the temptation to embed individual hostnames, user names, or one-off exceptions in automation code. The integration resolved facts; the customer’s response policy determined what those facts allowed.

Designing the analyst handoff

At the first review, the endpoint team asked us to show why an action was allowed before showing that it had run. The original work note emphasized connector status. It contained the GUID and request result but did not make the identity decision easy to challenge.

We reordered the handoff around evidence: incident hostname, resolved hostname, endpoint identifier, last-seen time, operating system, user or owner when known, identity result, agent state, asset role, exception, and proposed response. The action status followed those fields. We included a provider-console link tied to the resolved endpoint so that an analyst did not need to paste an opaque identifier into another system.

This changed the review conversation. Instead of asking only whether the connector had executed, the teams could ask whether the evidence supported the selected device and branch. On a mismatch, the two hostnames were visible together. On a stale result, the timestamp and policy outcome were explicit. On an approval path, the asset classification and exception reason explained why automation had stopped.

We avoided filling missing fields with convenient assumptions. An absent owner remained unavailable. An unknown role remained unknown. The note identified the enrichment source for material values so that analysts could distinguish incident evidence from provider evidence. This made holds actionable rather than presenting them as generic errors.

The same record supported shift handoff. Another analyst could understand what had been resolved, what remained uncertain, whether a request had been submitted, and who owned the next decision without rerunning the action or repeating all enrichment manually.

Making isolation stateful and verifiable

After identity and policy approved the action, the workflow checked current endpoint state before submitting isolation. An already isolated endpoint produced a completed/no-change result tied to the same endpoint identifier. A pending isolation request remained pending. Neither condition generated another request simply because a duplicate detection or impatient retry reached the workflow.

A successful API response meant the provider accepted the request; it did not by itself prove that containment had taken effect. We recorded the action as requested, retained its correlation data, and checked the endpoint’s resulting management state. Only provider confirmation allowed the ticket to describe the endpoint as isolated. A timeout or indeterminate response moved to reconciliation instead of blind resubmission.

Correlation was especially important when several findings referenced the same asset. The workflow associated the action with the verified endpoint and incident context, while idempotency checks prevented repeated requests from accumulating. A later run read current provider state first. It did not rely solely on the previous orchestration run’s status.

We treated release as a separate response action rather than a generic undo button. Before release, the workflow verified that the request referred to the same endpoint, that isolation was active, and that the approval remained valid under the customer’s process. It retained the original action correlation so an analyst could see which containment was being reversed.

After submitting release, the workflow again checked resulting provider state and agent health. An accepted request stayed requested until confirmation. If the device did not return to the expected managed state, the case remained open for endpoint-team review. Automatic containment did not imply automatic release, and success in one direction did not authorize the other.

Testing credible failures, not only the demo path

We developed the test set with analysts and endpoint engineers because they knew which ambiguous states appeared during real investigations. Before enabling the action path, we assigned an expected result to each case.

The positive case used a normal workstation whose incident and provider identities agreed, whose enrichment met freshness policy, whose agent was reachable, and whose classification allowed automatic response. We confirmed that the workflow displayed the evidence before the action, submitted one request, and verified the resulting isolation state.

The negative and exception cases included:

  • a valid GUID resolving to a different hostname;
  • a matching hostname with stale enrichment;
  • a missing hostname or incomplete device record;
  • an unavailable user map;
  • a correctly identified endpoint with an offline agent;
  • a shared jump host requiring approval;
  • a server or protected device classification;
  • an active maintenance exception;
  • duplicate detections while isolation was pending;
  • an endpoint already isolated before the run;
  • an accepted request followed by an ambiguous client response; and
  • release requested for a different or insufficiently correlated action.

For the critical test, we deliberately returned a plausible endpoint with a different hostname. The record had enough detail to look trustworthy, and the simulated agent was capable of receiving commands. The workflow stopped at the identity gate, recorded mismatch, displayed both names, and submitted no isolation request. It did not overwrite the provider hostname with the incident value or search for a more convenient result.

We also confirmed that stale and unreachable were not interchangeable. A stale record followed the customer’s confidence policy even if its last known agent state looked healthy. A fresh identity match with an offline agent remained the right device but could not yet be described as contained. Those paths assigned different owners and next steps.

Duplicate and timeout tests exercised the moments when an analyst was most likely to press retry. While the first isolation request was pending, another run reported pending. After provider confirmation, it reported already isolated. Following an ambiguous response, it reconciled against current state and correlation data before deciding whether any further action was safe.

Release tests covered authorization expiry, endpoint mismatch, already released state, provider delay, and unhealthy post-release management state. These cases prevented a convenient reversal path from bypassing the evidence and controls required for containment.

The delivered endpoint response

The final workflow was smaller than the first design because it stopped trying to infer its way around missing evidence. It resolved the supplied identifier, compared incident and provider identity, evaluated freshness and reachability separately, applied customer-owned asset policy, and only then allowed isolation. Every stop exposed a reason and an owner.

For public review material, we retained placeholders such as <host>, <user>, and <endpoint-id> rather than showing a live endpoint console. Real captures would have exposed device names, users, management links, and organizational details. The identity-gate diagram preserved the valuable part of the implementation: resolve, compare, classify, decide, and verify.

The customer’s analysts no longer had to reconstruct the same decision across several consoles for every eligible event. More importantly, the automation did not remove their ability to challenge the target. The ticket showed which endpoint it selected, how the evidence aligned, which exception policy applied, what request was made, and what state the provider confirmed.

The initial problem was not that the endpoint API returned no result. It returned a convincing wrong result. By working through that failure with the customer’s teams, we delivered a response that treated identity as a gate instead of a lookup side effect. The GUID remained useful, but it was one piece of evidence. Containment became defensible only when the rest of the context agreed.

ABOUT THE AUTHOR

Technoxi Security Engineering

Security Automation Team

We connect detection, case management, and response without hiding uncertainty.