Research · Defensive guidance

Detecting SPIFFE identity spoofing on a compromised Kubernetes node

Last reviewed:

Join cgroup or mount-namespace tampering, SPIRE Workload API issuance and protected-service authentication to expose a valid SVID held by the wrong process.

In this research
  1. Contribution
  2. The pattern
  3. Why it matters to cloud defenders
  4. ATT&CK mapping
  5. Detection guidance
  6. What to do now
Research snapshot
Type
Defensive guidance
Reviewed
2026-10-02
ATT&CK
T1611, T1528, T1649, T1550

Contribution

This post adds a detection method for SPIFFE/SPIRE identity spoofing that joins observation points the sources treat separately: cgroup or mount-namespace tampering on a Kubernetes node, a local request to the SPIRE Workload API, and use of the resulting SPIFFE ID from a place where the registered workload does not run. It covers two routes. Unit 42 demonstrated the first, which needs root on the node. A SPIRE maintainer change describes and closes the second, which did not need root and is not yet fixed in any published release. The contribution is defensive: it turns both into an identity-continuity tripwire that can tell a validly signed SVID apart from the process that was meant to receive it.

The pattern

SPIFFE replaces stored workload secrets with short-lived identities. SPIRE is its reference implementation. A workload asks the local SPIRE agent for an identity over the Workload API, normally exposed as a Unix domain socket. The agent identifies the calling process, gathers selectors through workload-attestor plugins, matches those selectors against cached registration entries, and returns an SVID for every matching entry. SPIFFE's implementation documentation describes that division of labour between the server, node agent and workload.

An SVID can be an X.509 certificate used for mutual TLS or a JWT used as a bearer token. Both are cryptographically valid after issue. That property is the point of SPIFFE, but it also explains the detection problem: a service receiving a wrongly issued SVID sees a credential signed by the expected trust domain. Certificate validation alone cannot tell whether the process presenting it was the process the agent attested.

Unit 42's 10 September 2026 SPIFFE/SPIRE identity misuse research demonstrates the break. The attack starts with root access to a Kubernetes node. The SPIRE Kubernetes workload attestor takes the caller's PID, reads its cgroup path from /proc/<pid>/cgroup (older releases could fall back to /proc/<pid>/mountinfo), extracts a pod UID and container ID, then asks the kubelet for the matching pod metadata. Namespace, service account, pod UID, image and related attributes become selectors.

Root can change the evidence the attestor reads. In the demonstration, the operator copies a victim workload's cgroup-shaped path into a new directory under /sys/fs/cgroup, writes the attacker's PID into that directory's cgroup.procs, and calls the local Workload API. Unit 42 notes that writing the PID into the victim's original cgroup would also have worked. The agent resolves the pod UID and container ID, finds the victim's registration entry, and returns its SVID to the attacker-controlled process. Unit 42's open-source Spooffe tool automates discovery of co-located workloads and collection of their identities.

The second route did not need root. SPIRE's container locator kept a fallback that read the caller's mount information when resolving cgroup membership, and a process controls its own mount view. SPIRE PR #7288 describes a process that creates a separate mount namespace, bind mounts a fabricated cgroup view, and steers the locator to an identity of its choosing. The change removes that forgeable fallback so container IDs come only from the kernel-controlled cgroup file. It merged on 17 September 2026 under the 1.16.0 milestone. The latest published release at the time of writing is 1.15.3, from 21 August 2026, which predates the fix. Verify the contents of the exact build you run rather than inferring protection from a milestone.

Neither route is a cryptographic forgery. The SPIRE agent returns a real credential because the local attestation evidence points at a registered workload. Neither is a remote pre-authentication flaw either: both need code execution on the node, and node integrity is part of SPIRE's trust model. The fix closes the unprivileged route only. It does not make a root-compromised node a trustworthy witness about its own cgroups. The useful defensive question is therefore narrower: once local evidence can be shaped, can the SOC see one process borrow another workload's identity before that identity reaches a protected service?

The answer cannot come from Kubernetes audit logs alone. Kubernetes auditing starts inside kube-apiserver and records API requests according to the configured policy. A local cgroup write, a bind mount and a Unix socket connection never cross that boundary. The SPIRE agent's pod metadata request can look like routine agent work because it is made with the agent's own service account. API audit remains useful for the entry path, such as creation of a privileged pod, but it does not record the selector substitution itself.

Why it matters to cloud defenders

Workload identity is often treated as the cure for secret theft. Short lifetimes remove static keys from images and environment variables, while automatic rotation reduces the value of material copied from disk. That is good engineering, but it does not remove the node from the trust chain. A compromised node sits below every pod scheduled on it and beside the local identity agent.

The area of impact follows placement, not namespace ownership. A low-value pod and a sensitive payment or deployment workload can be in different namespaces yet share a node. Once a process can reproduce the sensitive workload's selectors, namespace separation does not stop SVID issue. The victim pod does not need to be terminated, and its credential does not need to be read from its filesystem. Both the real workload and the impostor can hold valid material for the same SPIFFE ID at once, and expiry offers little comfort because the attacker asks the legitimate issuer for fresh material.

The impact also crosses into cloud accounts. A forged SVID can be accepted by a service mesh ingress, an internal API, a database proxy or an identity broker that exchanges SPIFFE identity for AWS, Azure or Google Cloud credentials. The manipulation happens on a Kubernetes node, but the useful access can sit in a cloud control plane.

That makes the first protected-service request more important than the SVID fetch by itself. Many legitimate workloads connect to the agent socket during start-up and rotation. SPIRE also supports multiple workload-attestor plugins and delegated identity patterns, so a blanket alert on Workload API access would be noisy. The suspicious condition is a break in continuity between four facts: the process that changed its cgroup membership or mount view, the process that requested an SVID, the selectors assigned to that identity, and the source that later presented it.

This pattern applies on managed Kubernetes as well as self-managed clusters. AWS, Azure and GCP can protect the worker-node boundary and retain control-plane logs, but none of those controls makes root inside a worker node trustworthy again. Provider audit logs may show how the node or cluster was changed. Runtime telemetry and service authentication logs show what happened after that change. The investigation must cross those planes, and the owners often differ: platform engineering holds SPIRE logs, Kubernetes operations holds node telemetry, and service owners hold mTLS peer records.

The practical consequence is uncomfortable but useful: treat root on a SPIRE agent node as compromise of every identity authorised on that node. Do not scope containment only to the pod where the first alert fired. Quarantine the node, identify every registration entry available to its agent, and review use of those SPIFFE IDs until the node and affected workload credentials have rotated through a trusted replacement.

ATT&CK mapping

The source demonstration begins after root compromise, so the entry technique depends on how that access was obtained. If a privileged container, sensitive host mount or runtime flaw allowed the attacker to leave a pod, the entry maps to T1611, Escape to Host. If root came through SSH, a vulnerable node service or stolen administrator credentials, use the technique supported by that evidence instead. Do not label every instance T1611 merely because Kubernetes is present. The unprivileged route needs no escape at all when the process already runs on the node.

The identity acquisition has two honest mappings. A harvested JWT-SVID is an application bearer token, which fits T1528, Steal Application Access Token. An X.509 SVID includes a certificate and private key, so its acquisition fits T1649, Steal or Forge Authentication Certificates. The attacker is causing the trusted issuer to mint material under the victim identity rather than breaking the signature algorithm.

Presentation of either SVID to a relying service is T1550, Use Alternate Authentication Material. This is where the chain reaches its operational payoff: the service accepts the SPIFFE ID and authorises an action. Record that action separately. A read from an internal secret service, an object store or a deployment API should carry its own service-specific technique once the logs prove it. The proof of concept stops at identity collection, so claiming data theft without a protected-service event would overstate the evidence.

Those mappings give analysts a clean boundary. T1611 is conditional entry evidence. T1528 or T1649 describes credential acquisition according to SVID type. T1550 describes use. The final service action establishes what the attacker gained. That sequence is more useful than attaching a single credential-access label to the whole incident.

Detection guidance

Start on the worker node. Runtime security or endpoint telemetry should retain successful file opens and writes under /sys/fs/cgroup, mount and bind-mount operations, unshare and setns calls affecting mount namespaces, Unix domain socket connections, process ancestry, login UID, container identity and node name. The first detector looks for an unexpected process opening a cgroup.procs file for writing. Container runtimes and service managers do this legitimately, so build the allowlist from observed executable paths rather than accepting any process running as root.

The following Falco-shaped rules show the two component signals. They are correlation inputs, not a claim that either event proves identity theft alone. Confirm the process allowlists and the SPIRE socket path before enabling an alert.

Code block YAML
- list: expected_cgroup_writers
  items: [systemd, kubelet, containerd, containerd-shim, runc, crun]

- list: expected_spire_clients
  items: [envoy, istio-proxy, spire-agent]

- rule: Unexpected process opens cgroup.procs for writing
  desc: Find a process outside the local runtime allowlist changing cgroup membership
  condition: >
    evt.is_open_write=true and
    fd.name startswith /sys/fs/cgroup/ and
    fd.name endswith cgroup.procs and
    not proc.name in (expected_cgroup_writers)
  output: >
    Unexpected cgroup membership write
    node=%evt.hostname process=%proc.name pid=%proc.pid
    parent=%proc.pname loginuid=%user.loginuid file=%fd.name
  priority: WARNING
  source: syscall

- rule: Unexpected process connects to SPIRE Workload API
  desc: Find a local client outside the workload allowlist connecting to the SPIRE agent
  condition: >
    evt.type=connect and
    fd.sockfamily=unix and
    fd.name contains /run/spire/sockets/agent.sock and
    not proc.name in (expected_spire_clients)
  output: >
    Unexpected SPIRE Workload API client
    node=%evt.hostname process=%proc.name pid=%proc.pid
    parent=%proc.pname loginuid=%user.loginuid socket=%fd.name
  priority: NOTICE
  source: syscall

The socket path above is the SPIRE quickstart default. The SPIRE Helm chart uses /run/spire/agent-sockets/spire-agent.sock, so use the socket_path your agents are configured with. Falco's supported-fields reference documents evt.is_open_write, process fields, file names and socket-family fields for syscall rules. Falco maintainers also advise that Unix socket paths are available on the connect event through fd.name, rather than on the earlier socket creation event. Test this on the deployed Falco driver because captured fields can vary by driver and kernel.

Join the two alerts by node and user or process ancestry in a short window. The exact duration should follow normal node operations; two minutes is a useful starting hypothesis, not a universal constant. Raise severity when the value written to cgroup.procs moves a process into a cgroup owned by a different pod UID. That comparison needs a pod inventory captured at event time, because pod cgroups disappear quickly. A process changing membership within its own pod is less interesting than one crossing into another workload's cgroup immediately before opening the Workload API socket. Shells, tee, unfamiliar binaries and a copied spire-agent command-line client also raise severity.

For the unprivileged route, look for a short sequence from one process tree: create a mount namespace, bind mount a cgroup-shaped file or directory over /proc or /sys/fs/cgroup, then connect to the agent socket. Development tools and service managers can create namespaces, which is why the socket connection and the cgroup-shaped mount target belong in the same condition. On a build that contains the fix, the same attempt should fail attestation instead of returning an SVID. Keep the node event even when no identity follows. Repeated failures show somebody testing the boundary, and a patch that blocks issuance without an alert gives an attacker unlimited quiet attempts against nodes that missed the update.

The SPIRE agent is the second evidence source. With JSON logging at debug level, the Workload API handler records Fetched X.509 SVID and Fetched JWT SVID with the caller PID, the issued SPIFFE ID, whether a registration matched, and the SVID lifetime. Failed requests produce records such as Workload attestation failed or No identity issued, the latter with registered=false and the attested selectors when log_selectors is configured. Debug output is voluminous and selector values can be sensitive, so run it on a pilot set of nodes with an owner rather than across every node indefinitely. Aggregate metrics are not a substitute. SPIRE supports Prometheus, StatsD, DogStatsD and M3 collectors, according to its telemetry configuration, and request-rate changes can guide triage, but counters carry no PID, selector or source-address evidence. A burst of identity requests may also come from a rollout or node recovery.

Resolve the PID independently. At SVID issue time, ask the container runtime and the Kubernetes API which pod the PID belongs to, then compare that pod's node, UID, namespace and service account with the registration selectors for the issued SPIFFE ID. Do not reuse the cgroup-derived answer as its own check, because that is the evidence the attacker shaped. The detector is strongest when runtime metadata says the process belongs to pod A while SPIRE issued an identity registered to pod B.

The final event comes from the relying service or service mesh. Log the authenticated peer SPIFFE ID, source IP, destination, requested operation and decision. For mTLS, capture the peer URI SAN rather than the full certificate. For JWT-SVIDs, record the validated subject and audience, never the token. Resolve the SPIFFE ID back to its registration entry and expected pod selectors, then compare the network source with the live pod inventory. A high-confidence alert has all of these properties:

  1. An unexpected process writes a PID to cgroup.procs, or presents a new mount-namespace view, on a node.
  2. A process in the same execution context connects to the SPIRE agent socket.
  3. SPIRE issues an identity that independent runtime metadata says does not belong to that process.
  4. A protected service accepts that SPIFFE ID while the request source is the node address, an unknown workload address, or a pod that does not satisfy the identity's selectors.
  5. The accepted request performs an operation the source workload has not used in its established baseline.

Treat the first three as a high-confidence attempted spoof and add the last two for confirmed use. If SPIRE logs are unavailable, the downstream location mismatch can still catch abuse, but it loses the direct link to issuance. If service logs are unavailable, node tampering plus a mismatched issuance still deserves containment.

False positives cluster around node maintenance. Runtime upgrades, pod churn, privileged diagnostics and incident-response tooling can change cgroups, enter namespaces or call the Workload API. Service meshes can route a connection through a proxy whose source address differs from the application pod, so tune against pod UID and node placement rather than IP alone, and log both the proxy and the original workload identity where the protocol allows. Suppress only when a change ticket, known executable path and expected service account agree. A broad exclusion for UID 0 destroys the detector because root is the prerequisite for the demonstrated route.

What to do now

  1. Inventory SPIRE agent versions and build provenance by node. Check whether the deployed artefact contains the PR #7288 change; do not assume it from the 1.16.0 milestone. Treat nodes on 1.15.3 or earlier, or on custom builds, as exposed to the unprivileged route until verified, and roll the first release that includes the fix through staging before production.
  2. Inventory each SPIRE agent's reachable registration entries. Treat the set as the node's identity area of impact, then check whether sensitive and low-trust workloads share that node.
  3. Review exposure of the Workload API socket. Remove hostPath mounts that hand it to unrelated pods, and inspect privileged workloads that can reach both the socket and host cgroups. This will not stop root, but it removes accidental paths and makes socket access more meaningful in telemetry.
  4. Collect runtime events for cgroup.procs writes, mount-namespace changes, bind mounts and Unix connect calls to the configured socket. Verify that node name, PID, process ancestry, login UID and file or socket path reach the SIEM, and pilot SPIRE debug logging on a small node set.
  5. Add peer SPIFFE ID and network source to service-mesh or application authentication logs. Without both fields, a valid spoofed SVID is almost indistinguishable from its intended holder at the service boundary.
  6. Test the component alerts during an authorised exercise. Unit 42 released Spooffe for defensive assessment, but run it only on an isolated or approved node because it attempts to retrieve co-located identities. Replay the unprivileged route against a candidate build in a non-production cluster and confirm it fails and alerts.
  7. On a node-level alert, cordon and isolate the node, preserve runtime and audit evidence, replace it from trusted infrastructure, rotate affected workload identities, and review every protected-service action made under identities authorised to that node. If a broker exchanged the SVID for AWS, Azure or Google Cloud credentials, review the cloud audit trail through token expiry and revoke sessions where the platform permits it.

The trust decision is the thing to preserve during triage. Kubernetes API audit explains control-plane activity, runtime events and SPIRE logs show manipulation of local attestation evidence, and service logs show whether the borrowed identity achieved anything. Joined together, they expose the gap that signature validation cannot: a credential can be genuine while its holder is not.

01 ATT&CK references