NEWS

Cisco MINT Partner! Learn more →

Automation & AI
2026-10-11
11 min read

Teaching Agents eBPF: Why We Built a Cilium MCP Server for Kubernetes Triage

Giving AI coding agents raw bash access to debug Kubernetes clusters leads to context window flooding, hallucinated commands, and risky mutations. Here is how we built a Model Context Protocol (MCP) server for Cilium and Hubble that gives agents structured, read-only eBPF network telemetry, strict output contracts, and guarded dry-run actions.

Kubernetes
Cilium
eBPF
Model Context Protocol
MCP
Hubble
Network Security
AI Agents

During an incident triage call at 02:00 AM, a payment ingestion microservice in a production Kubernetes cluster began dropping 15% of outbound API requests to an internal fraud scoring backend.

The platform engineer on call did what many engineers are experimenting with today: they launched an AI coding assistant in their terminal, gave it shell access via a terminal hook, and typed: "Figure out why pods in namespace payments cannot reach fraud-api.internal.svc."

What followed was an operational disaster.

First, the agent executed an unconstrained kubectl logs command across twenty pods without line limits, dumping 42 megabytes of raw JSON logs into its context window. Having blown through its token budget, the agent truncated its memory, forgot earlier hypotheses, and hallucinated a non-existent cilium-dbg netpol --all command. Frustrated by the shell error, it decided to take matters into its own hands and executed kubectl delete pod -l app=payments-api—restarting the pods simultaneously, dropping active user transactions, and wiping the transient socket states needed to diagnose the issue.

The problem was never the microservice code. It was an eBPF network policy rule update deployed twenty minutes earlier by another team. Cilium was silently dropping SYN packets at the socket layer because a namespace label had changed.

This incident crystallized an uncomfortable reality: giving AI agents raw bash access to a Kubernetes cluster during an outage is a liability. Unbounded shells produce context pollution, hallucinated syntax, and unvetted destructive actions.

To make AI triage viable in production, we built technoxi/cilium-mcp—a dedicated Model Context Protocol (MCP) server that provides AI assistants with structured, read-only Cilium and Hubble telemetry, strict output contracts, and guarded action primitives.


The Failure Modes of the Unstructured Shell

When engineers connect Large Language Models (LLMs) to infrastructure, they usually default to an agent with a bash tool. While bash is flexible, it fails in production for four distinct reasons:

  1. Context Window Flooding: A single kubectl get pods -A -o yaml or raw Hubble flow stream can emit tens of thousands of tokens. Once an agent's context fills with low-signal JSON, its reasoning degradation accelerates, leading to forgotten constraints and hallucinations.
  2. Fragile Text Parsing: When an agent runs cilium endpoint list, it receives an ASCII table formatted for human eyes. The LLM must spend precious attention tokens attempting to regex column alignments and parse truncated strings, introducing subtle parsing errors.
  3. No Execution Guardrails: An agent with a generic shell does not distinguish between a benign inspection command (kubectl get cnp) and a catastrophic mutation (kubectl delete -f policy.yaml). Even if prompted to "be careful," an autonomous reasoning loop under pressure will attempt aggressive changes if it believes they solve the prompt.
  4. Zero Auditability: Shell history (.bash_history) is unstructured. It does not record the agent's intent, the evidence it weighed, or why it chose a particular command.

To bridge AI agents to infrastructure safely, we do not need bigger context windows. We need typed, contract-driven tool adapters.


Why Model Context Protocol (MCP) Fits Kubernetes Networking

Anthropic's Model Context Protocol (MCP) provides a standard open specification for connecting AI applications to external data sources and tools. Instead of exposing raw system calls, an MCP server exposes discrete functions defined by JSON Schema contracts.

In cilium-mcp, the LLM interacts exclusively with strongly typed tools. The agent never executes a raw subshell; it calls high-signal analytical tools and receives deterministic, curated JSON envelopes:

[AI Assistant / Coding Agent (Claude, Cursor, Antigravity)]
                    │
                    ▼ (JSON-RPC 2.0 over Stdio or HTTP /mcp)
       ┌──────────────────────────────┐
       │      cilium-mcp-server       │
       └──────────────────────────────┘
                    │
         ┌──────────┴──────────┐
         ▼                     ▼
┌──────────────────┐  ┌──────────────────┐
│ Kubernetes Client│  │   Cilium / Hubble │
│ (kube-apiserver) │  │   (eBPF Observer) │
└──────────────────┘  └──────────────────┘

The server supports two standard transport layers:

  • Stdio (Standard I/O): For direct integration with local CLI agents, IDEs, and developer workflows.
  • Streamable HTTP (/mcp): For remote, headless platform agent teams running in containerized orchestration environments.

Tool Families: From Cluster Context to Root Cause Analysis

We organized the MCP server into three distinct tool families designed to mirror how experienced Site Reliability Engineers (SREs) actually triage an incident.

1. Context Tools (High-Signal First)

Rather than making an agent guess where to look, context tools aggregate multi-dimensional cluster telemetry into dense, high-signal summaries:

  • context.cluster_summary: High-level nodes, Cilium daemonset health, and controller statuses.
  • context.workload_health: Identifies crashing pods, OOMKilled containers, and unready endpoints in a target namespace.
  • context.network_insights: Summarizes recent Cilium eBPF drops, DNS lookup latencies, and cross-node encapsulation issues.
  • context.rca_pack: The ultimate triage tool. Gathers workload logs, events, active network policies, and recent Hubble packet drops into a single cohesive incident dossier.

2. Investigation Tools (Deep-Dive Analysis)

When an agent identifies an anomaly, it pivots to targeted diagnostic probes:

  • investigate.pod: Inspects container specs, restart counts, and Cilium endpoint identity metadata.
  • investigate.service_path: Traces the complete data path from client pod through Kubernetes Service VIP, kube-proxy replacement (eBPF map), to backend pod endpoints.
  • investigate.policy_impact: Simulates and evaluates whether a specific CiliumNetworkPolicy (CNP) or CiliumClusterwideNetworkPolicy (CCNP) selects an endpoint and blocks traffic.
  • investigate.flow_query: Queries the Hubble flow buffer for real-time L3, L4, and L7 packet decisions with fine-grained filters (verdict: DROPPED, source pod, destination port).

3. Guarded Actions (Dry-Run First)

When remediation is necessary, actions are strictly constrained:

  • act.rollout_restart: Safely restarts a deployment or daemonset.
  • act.scale: Adjusts workload replica counts within strict bounds.

The Output Contract: Eliminating Hallucinations

A major breakthrough in building cilium-mcp was establishing a mandatory output envelope. Every tool in the server—whether checking cluster health or querying eBPF drops—returns data formatted according to a shared JSON schema:

{
  "executive_summary": "Egress packets from 'payments-api' to 'fraud-scoring' are being dropped by Cilium eBPF policy engine.",
  "top_findings": [
    {
      "severity": "HIGH",
      "confidence": 0.98,
      "title": "CiliumNetworkPolicy 'deny-unlabeled-egress' dropping traffic",
      "description": "Destination pod lacks label 'env=prod', causing selector mismatch in egress rule #2."
    }
  ],
  "evidence": {
    "dropped_flows_count": 482,
    "drop_reason": "Policy denied (Identity 1042 -> 3105)",
    "policy_name": "deny-unlabeled-egress",
    "namespace": "payments"
  },
  "causal_graph": [
    "Namespace label 'env=prod' removed in commit a109f",
    "Cilium endpoint identity updated from 3104 to 3105",
    "Egress rule selector failed to match identity 3105",
    "eBPF tail-call dropped SYN packets at socket layer"
  ],
  "actions": [
    {
      "action_type": "patch_label",
      "target": "pod/fraud-scoring-78b9-x21z",
      "patch": "{\"metadata\":{\"labels\":{\"env\":\"prod\"}}}"
    }
  ],
  "commands": [
    "hubble observe --namespace payments --verdict DROPPED --to-pod fraud-scoring",
    "cilium-dbg endpoint get 1042"
  ],
  "rollback_and_verification": {
    "verification_command": "hubble observe --namespace payments --follow",
    "expected_result": "Verdict: FORWARDED"
  },
  "meta": {
    "execution_time_ms": 142,
    "version": "1.0.0"
  }
}

This envelope accomplishes three things:

  1. Ranked Findings: The agent receives prioritized findings with explicit confidence scores, preventing it from rabbit-holing down irrelevant warnings.
  2. Causal Graphing: Forces the model to understand the dependency chain before suggesting changes.
  3. Reproducible Human Verification: The commands array gives the human engineer exact, copy-pasteable CLI commands to verify the agent's claims independently in their own shell.

Deep Dive: Pinpointing eBPF Drops with Hubble

In standard Kubernetes networking without Cilium, tracking down why two pods cannot talk often requires tcpdump, packet captures across overlay interfaces, or inspecting complex iptables chains across worker nodes.

With Cilium, packet filtering happens inside the Linux kernel via eBPF programs attached to traffic control (tc) hooks and socket layer (sockops). When a packet is dropped, Cilium's eBPF datapath emits an event containing the source security identity, destination security identity, and the exact drop reason.

Inside investigate.flow_query, the MCP server queries the local or remote Hubble gRPC service. When an agent queries for drops, it receives structured flow records:

{
  "time": "2026-10-11T02:14:32.189Z",
  "verdict": "DROPPED",
  "drop_reason_desc": "Policy denied",
  "drop_reason": 133,
  "source": {
    "identity": 1042,
    "namespace": "payments",
    "pod_name": "payments-api-54d9c74846-9zq8l",
    "labels": ["app=payments-api", "tier=backend"]
  },
  "destination": {
    "identity": 3105,
    "namespace": "payments",
    "pod_name": "fraud-scoring-68db6554b7-jkm2x",
    "labels": ["app=fraud-scoring"]
  },
  "traffic": {
    "protocol": "TCP",
    "source_port": 49182,
    "destination_port": 8443
  }
}

Because the output is structured, the AI assistant can immediately compare the destination labels (["app=fraud-scoring"]) against the active CiliumNetworkPolicy YAML manifest:

apiVersion: "cilium.io/v2"
kind: CiliumNetworkPolicy
metadata:
  name: "allow-fraud-scoring"
  namespace: "payments"
spec:
  endpointSelector:
    matchLabels:
      app: payments-api
  egress:
  - toEndpoints:
    - matchLabels:
        app: fraud-scoring
        env: prod       # <--- The missing label causing drop reason 133

Within two tool calls, the agent identifies that the destination pod is missing env=prod. It did not need to run shell commands, grep through logs, or guess at network topologies.


Action Guardrails: Preventing Production Accidents

Allowing an AI agent to suggest changes is helpful; allowing it to execute them requires strict controls. In cilium-mcp, action tools (act.*) enforce three layers of programmatic defense:

1. Mandatory Dry-Run First

Every action tool requires a dry_run: boolean parameter. If an agent calls an action with dry_run: true (or omits the parameter), the server executes the mutation using Kubernetes server-side dry run (--dry-run=server). The server returns the simulated outcome, allowing the agent to confirm that its patch is syntactically and semantically valid before applying it.

2. Approval Tokens for Live Mutations

To prevent autonomous agents from triggering live changes without human authorization, the server can be started with an approval token:

export CILIUM_MCP_APPROVAL_TOKEN='sec-auth-98421-prod'

If an agent attempts to call act.rollout_restart with dry_run: false without passing the matching token, the call is rejected with a 403 Forbidden error. In an automated pipeline, a human operator reviews the agent's proposed plan, approves the token, and unlocks the execution.

3. Blast-Radius Clamps

To prevent scaling accidents, the server enforces hard ceiling limits:

export CILIUM_MCP_MAX_SCALE_DELTA=5

If an agent requests scaling a deployment from 3 replicas to 50 replicas, the server rejects the request immediately, preventing compute exhaustion and cloud bill shocks.

4. Immutable JSONL Audit Trail

Every successful action is appended to a local audit log:

export CILIUM_MCP_AUDIT_LOG_PATH='/var/log/cilium-mcp-audit.jsonl'

The log records the exact tool called, the agent's input arguments, the target resource, the timestamp, and the user identity that authorized the token.


Getting Started: Running cilium-mcp Locally

You can spin up the server locally to test against a development cluster or kind/minikube environment.

1. Installation

Clone the repository and install dependencies inside a virtual environment:

git clone https://github.com/technoxi/cilium-mcp.git
cd cilium-mcp
python3 -m venv .venv
source .venv/bin/activate
pip install -e .

2. Run Modes

For standard local execution (e.g., configuring Claude Desktop or Cursor):

cilium-mcp-server

To run as an HTTP microservice for agent teams:

cilium-mcp-server --transport http --host 127.0.0.1 --port 8000 --path /mcp

3. Direct CLI Testing (No GUI Required)

The repository includes a standalone test CLI in scripts/mcp_cli_test.py to verify tool calls directly:

# List all registered MCP tools
python3 scripts/mcp_cli_test.py --url http://127.0.0.1:8000/mcp list-tools

# Run a namespace inventory query
python3 scripts/mcp_cli_test.py --url http://127.0.0.1:8000/mcp call \
  --tool context.namespace_inventory \
  --args '{"namespace":"payments"}'

# Execute a comprehensive Root Cause Analysis pack
python3 scripts/mcp_cli_test.py --url http://127.0.0.1:8000/mcp call \
  --tool context.rca_pack \
  --args '{"namespace":"payments","workload":"payments-api","since_minutes":10}'

Engineering Takeaways

The transition from human CLI debugging to AI-assisted infrastructure operations requires rethinking how tools are exposed.

  1. Structured lenses beat raw shells: LLMs make significantly better decisions when presented with curated, typed telemetry envelopes rather than raw command-line text.
  2. eBPF data is ideal for AI analysis: Because Cilium captures deterministic network identities and drop reasons at the kernel level, eBPF telemetry provides the exact ground truth AI agents need to diagnose complex distributed failures.
  3. Guardrails must be structural, not conversational: You cannot prompt an agent into being safe. Safety requires dry-run enforcement, parameter delta clamps, and cryptographic approval tokens built directly into the tool layer.

ABOUT THE AUTHOR

Technoxi Security Engineering

Security Automation Team

We build automation that connects detection, change management, and response — without cutting the corners that keep security teams in control.