NEWS

Cisco MINT Partner! Learn more →

Automation & AI
2026-09-15
9 min read

Proving the Message Was Quarantined

Broad remediation feels decisive and cleans the inbox. Then someone asks which messages you acted on, and the answer is 'we ran a search'. Targeted message-ID remediation answers that question per mailbox, per message.

Email Security
SOAR
Phishing Remediation
Microsoft Graph
Incident Response
Security Automation
Email identifier state machine: track evidence and provider message identifiers through verified remediation.

A customer asked our joint engineering team to automate a familiar incident-response task: when an analyst confirmed a phishing message, remove every delivered copy and write the result back to the incident and ServiceNow. The existing manual process searched by sender and subject, then moved everything that looked related. It was fast, but the customer’s requirement was more exacting. The automation had to show which message in which mailbox it had touched, preserve enough evidence for review, and never turn an uncertain API response into a success.

That last requirement changed the project. We were not building a convenient search-and-delete button. We were building a per-message accounting system whose containment action happened to be quarantine.

The first successful run that was not a success

Our early proof of concept looked healthy. It accepted evidence from an XDR incident, submitted a move request, received an HTTP success response, and posted “phishing email quarantined” to the ticket. During review, the customer team picked one result and tried to locate it by the recorded identifier. They could not.

The run had mixed several meanings of “message ID.” The incident supplied an alert evidence ID, the message header contained an RFC 5322 Message-ID, the email-security platform had its own provider message ID, and the mailbox API addressed an item within a particular mailbox. The move response added a remediation record ID. All were legitimate identifiers, but none were interchangeable.

We also learned that acceptance was not completion. A provider could accept a batch while individual messages remained pending or failed. A mailbox move could change an ordinary Outlook item ID. One message could exist in several recipient mailboxes under separate provider objects. The green run status proved only that our orchestration had reached its last step.

Together, the customer’s responders and our automation engineers replaced the single “quarantined” outcome with a lifecycle:

identified
  -> validated
  -> remediation requested
  -> remediation accepted
  -> verification pending
  -> verified

Terminal outcomes included already_quarantined, not_found, invalid_identifier, permission_denied, rate_limited, failed, and pending_verification. That vocabulary was deliberately unglamorous. It let an analyst distinguish reduced exposure from proven completion.

Following each identifier to its owner

We traced every field from the source incident through the provider documentation and actual responses. The resulting identifier map became part of every run rather than tribal knowledge:

IdentifierOwner and purposeProject decision
RFC 5322 Message-IDSending system; useful across headers, alerts, and exportsPreserve as evidence, never assume it is the remediation key
Provider message IDEmail-security provider’s search, status, and action APIsUse as the primary action target
Alert evidence IDDetection/XDR incident recordKeep as the reason the target was selected
Mailbox item IDA message resource within one mailboxAlways retain mailbox scope; account for ID changes after moves
Conversation IDA thread containing potentially different messagesUse for investigation, not precise remediation
Remediation IDThe provider’s action or operationUse to reconcile and verify the request, not to identify the message

Microsoft documents that ordinary Outlook resource IDs can change when an item moves. Immutable IDs remain stable for an item within the same mailbox only when Prefer: IdType="ImmutableId" is used consistently; they are not universal email identifiers (Microsoft immutable IDs). Microsoft’s move operation creates a new copy in the destination and removes the original, explaining why a GET with an old ordinary ID can return 404 after a successful move (Microsoft message move).

The email-security provider had a different contract. Cisco ETD’s workflow consumes email message IDs and exposes search, remediation, and status operations; its provider ID is the action key, not an XDR evidence ID (Cisco XDR workflow). We therefore carried both provider and mailbox references without pretending one platform’s behavior applied to the other.

A normalized target looked like this:

{
  "source_incident_id": "<incident-id>",
  "alert_evidence_id": "<evidence-id>",
  "rfc5322_message_id": "<message-id@example>",
  "provider": "<email-security-provider>",
  "provider_message_id": "<provider-message-id>",
  "mailbox": "<[email protected]>",
  "requested_action": "quarantine",
  "selection_reason": "confirmed malicious attachment"
}

This also settled a design debate about discovery. Sender, subject, URL, and time-window searches remained useful for finding candidates, but search results could not silently become the action scope. The workflow validated the provider ID and mailbox, froze the target list, and only then submitted exact IDs. A missing provider ID stopped that item; it never triggered a fallback such as “delete everything from this sender.”

Designing the containment and accounting paths together

We separated the implementation into discovery, normalization, validation, authentication, remediation, verification, and reporting. That separation made partial failure explicit. A malformed record could be rejected while valid records continued, but the campaign roll-up remained partial.

Deduplication used (provider, mailbox, provider_message_id). Deduplicating on the RFC header would have collapsed delivered copies in different mailboxes into one target. Deduplicating on subject would have been worse: a legitimate invoice and a malicious reply could share one subject or conversation.

The team chose quarantine as the default reversible action. Junk or deleted items did not satisfy the customer’s containment policy, and hard delete created too much blast radius while identifier handling was still being proven. Release and reclassification became separate, approved actions with their own correlation records. Microsoft Graph exposes several remediation actions, including junk, deleted items, soft delete, hard delete, and inbox; recording the requested action exactly avoided reporting all of them as quarantine (Microsoft analyzedEmail remediation).

Cisco ETD permits up to 100 messages in a remediation request, but batching did not change our unit of accountability (Cisco remediation API). For each batch we retained its ID, ordered targets, timestamp, provider response, and a separate result for every message. We used smaller batches when selection was ambiguous or responses were easier to isolate one at a time.

The action record and the verification record stayed distinct:

{
  "provider_message_id": "<provider-message-id>",
  "mailbox": "<[email protected]>",
  "requested_action": "quarantine",
  "request_state": "accepted",
  "remediation_id": "<action-id>",
  "verification_state": "pending",
  "verified_folder": null,
  "last_checked_at": "<timestamp>"
}

After submission, a worker waited according to provider behavior and queried status. Cisco’s Status API exposes recent actions, including action, folder, initiator, and status, and notes that remediation can require time before checking (Cisco ETD Status API). We matched the current request by remediation ID, timestamp, initiator, or correlation fields instead of assuming the newest history entry belonged to us. That mattered when a user or another responder moved the message at nearly the same time.

Targeted remediation loop: evidence produces message-recipient pairs, the remediate call returns 202, each pair is re-queried for its final state, failures are retried once and then logged for a human.

Retries were designed around ambiguity. A network failure before transmission could retry with bounded exponential backoff. A timeout after transmission first reconciled by action and message ID, because resubmitting might duplicate a completed move. A 401 refreshed the token once; 403 stopped for permission review; 404 triggered an identifier and mailbox-scope check, never a broader search. For 429, the worker honored Retry-After and reduced concurrency. Cisco’s published ETD limits describe 2 requests per second per tenant, a burst of 4, and 10,000 daily requests, although we treated current provider documentation as authoritative because limits can change (Cisco ETD rate limiting).

Failure tests became the acceptance test

The happy path had passed on our first prototype, so the joint test plan concentrated on ways it could lie. We submitted a valid RFC header where a provider ID was required and expected invalid_identifier, not a broad lookup. We delivered copies with the same RFC header to multiple mailboxes and confirmed that all scoped targets survived normalization. We passed a conversation ID accidentally, duplicated one provider target, and tested a message already in quarantine.

Then we exercised timing and operational failures: one item failed inside an accepted batch; status remained pending beyond the run deadline; a timeout occurred after the provider may have accepted the action; a token expired during a batch; the service returned 429 and 5xx; ticket creation failed after remediation succeeded. We also moved a Graph-backed message and verified that the old ordinary item ID could disappear without changing the provider action result.

For each case we checked four surfaces: provider state, workflow record, source incident, and ServiceNow. The most important assertion was negative: no uncertain state could render as “quarantined.” A timed-out verification became pending_verification; a user’s prior move became already_quarantined with the observed initiator; one rejected mailbox made the campaign partially completed.

We also replayed the same incident and target after a completed run. The workflow checked current state before issuing another move, preserved the original remediation reference, and reported an idempotent outcome rather than manufacturing a second success. For a confirmed false positive, a separately approved release referenced the original quarantine action and appended a correction instead of editing history. Those tests mattered because production incidents are rarely linear: analysts retry buttons, users move mail, and verdicts change after more evidence arrives.

The analyst-facing note was intentionally concise:

-- Email remediation summary
Source incident: <incident-id>
Overall result: partially completed

-- Messages
- <provider-message-id-a> / <mailbox-a>: quarantine verified
- <provider-message-id-b> / <mailbox-b>: already quarantined by user action
- <provider-message-id-c> / <mailbox-c>: verification pending
- <provider-message-id-d> / <mailbox-d>: invalid provider ID; no action sent

The detailed run record retained request and verification timestamps, pre- and post-move references, retries, throttling events, final folder, and failure reason. It deliberately excluded tokens, secrets, and authorization headers. If ServiceNow failed, the source incident still received the remediation outcome and the ticket update entered reconciliation; reporting failure never erased containment evidence.

The customer outcome

By the end of the project, responders still initiated remediation from the incident, but they no longer had to trust a broad search or a green orchestration icon. Each delivered copy had a scoped identity, an action record, and a verified final state. Reviewers could distinguish a provider-accepted request from a completed move, and analysts could release a false positive without rewriting the original history.

The workflow was slightly more conservative than the prototype. It sometimes finished with “pending verification” or “partial” where the old process would have announced success. That was the desired result: uncertainty became visible work instead of hidden risk.

Our shared lesson was that targeted quarantine is not primarily an API-call problem. It is an identity-and-state problem. Search discovers candidates; provider IDs and mailbox scope define targets; remediation IDs track actions; status proves outcomes. Once those roles were explicit, batching, retries, tickets, and audit evidence all became simpler—and the customer could answer the question that started the project: exactly which messages did we act on, and what happened to each one?

ABOUT THE AUTHOR

Technoxi Security Engineering

Email Security Automation

We build phishing remediation that can prove, after the fact, exactly which messages it touched and what state each one ended in.