MCP Error Fingerprinting
MCP error fingerprinting: opentel-mcp groups identical CallToolResult.isError and thrown failures into one SHA-256 ID — cardinality-safe by construction.
MCP error fingerprinting is what turns a flood of unrelated-looking
failures into groups: opentel-mcp's Deep Failure Fingerprinting,
shipped in v0.4.0, computes a stable 16-character identifier for every
thrown error and every CallToolResult.isError: true tool-level
failure, attached as span attributes alongside the span status that
Silent Failures already marks ERROR. Ten
calls that fail the same underlying way — even with different user
IDs, timestamps, or UUIDs in the error message — group under one
fingerprint instead of showing up as ten unrelated errors. It runs
synchronously, locally, with no network call, and never throws.
TL;DR
Every failure gets a mcp.failure.fingerprint (16 hex chars, SHA-256
truncated), a human-readable mcp.failure.signature
(errorClass@fn:line), and one of 8 mcp.failure.category values.
Fingerprinting is on by default (fingerprinting: true) and can't
throw — if anything about it fails internally, it falls back to a
fixed identity (0000000000000000) rather than breaking the tool call
it's instrumenting.
How is the fingerprint computed?
computeFingerprint() hashes a small set of normalized inputs — not the
raw error message or raw stack trace, which would defeat grouping the
moment a UUID or timestamp changed between two otherwise-identical
failures:
v1|<errorClass>|<category>|<origin>|<toolName>|<normalizedMessage>|<stackSignature>errorClass is the error's constructor name ("TypeError",
"ZodError", or "MCPToolError" for a tool-level isError: true result
with no thrown exception behind it). category is one of the 8 buckets
below. origin is tool_error, thrown, or transport. The message
and stack trace go through a normalization pass first — stripping
UUIDs, timestamps, and similar high-cardinality noise — before being
folded into the hash, which is exactly what makes two failures with
different incidental details still hash identically.
That versioned string is hashed with SHA-256 and truncated to the first 16 hex characters:
// src/fingerprint/hash.js
export function hashInputs(input) {
return createHash('sha256').update(input, 'utf8').digest('hex').slice(0, 16);
}The package's own reasoning for that specific choice, from the source:
SHA-256 is native to Node's crypto module and fast enough to stay well
inside a p99 < 200µs budget for the whole pipeline; 64 bits of hash gives
roughly 1-in-4-billion collision odds at 100,000 unique fingerprints,
which is enough for grouping errors, not a security property; truncating
a cryptographic hash for identity purposes (not security) is a
recognized practice; and adding a dependency like xxhash or murmurhash
for a small speed gain wasn't worth it for a package that otherwise has
none.
What are the 8 failure categories?
Classification runs through seven ordered classifiers — the first one
that matches wins, and internal is a catch-all that always matches, so
it has to run last:
validation— bad input (Zod/Joi/Yup-shaped errors, "invalid"/"required" wording)timeout— an operation timed out (TimeoutError,ETIMEDOUT, ...)network— a connection failed (ECONNREFUSED,FetchError, ...)auth— 401/403, "unauthorized"/"forbidden" wording, plus (v0.9.0+) "permission denied", "access denied", "not authorized", "insufficient permission(s)", andEACCES/EPERMdependency— a downstream service or package failed (a database driver, an SDK, ...)serialization— malformed JSON, "unexpected token" wordinginternal— nothing more specific matched; the catch-all
Fingerprint values changed on v0.9.0 upgrade for auth-classified messages
The auth classifier missed "permission denied" and "access denied" —
the standard OS/CLI/IAM phrasing for a permission failure — before
v0.9.0; it only recognized HTTP-status-derived wording (unauthorized,
forbidden, authenticat(e|ion)), 401/403 status codes, and a handful
of known auth-library error names. Messages using that OS/CLI phrasing
fell through every classifier and landed in internal instead. As of
v0.9.0, it also matches "not authorized", "permission(s) denied",
"access denied", "insufficient permission(s)", and Node's own
EACCES/EPERM error codes — still not bare "authorized" or
"permission" alone, to avoid false positives like "user denied the
permission request".
Because category is one of the hashed inputs computeFingerprint()
combines into mcp.failure.fingerprint, this changes the fingerprint
for affected messages. A permission-denied failure that previously
classified as internal now classifies as auth — the fields feeding
the hash change, so the fingerprint itself changes. If you alert or
dashboard on a specific mcp.failure.fingerprint value for a
permission error, expect a new value after upgrading to v0.9.0. This
does not amend the closed 8-category taxonomy below — auth
already existed; this is pattern coverage for when it fires, not a new
category.
An 8th value, unknown, isn't one of the seven classifiers — it's what
computeFingerprint() itself falls back to if the whole pipeline hits an
internal error (see "What happens if fingerprinting itself fails?"
below), or what a custom classifier list falls back to if it omits
internal.
Order matters: a timeout error whose message happens to mention
"network" still classifies as timeout, because the timeout
classifier runs before the network one.
What does a fingerprinted span look like?
Illustrative
Real attribute names and values from src/fingerprint/attributes.js,
shaped the way a trace viewer would display them — not a literal
capture from a running instance.
{
"mcp.failure.fingerprint": "a3f4c8e2b1d09f77",
"mcp.failure.signature": "TimeoutError@fetchUpstream:42",
"mcp.failure.category": "timeout",
"mcp.failure.origin": "tool_error",
"mcp.failure.error_class": "TimeoutError"
}mcp.failure.signature is built as `${errorClass}@${topFrame.fn}:${topFrame.line}`,
truncated to 60 characters — human-readable enough to recognize at a
glance, without embedding the full stack.
Why don't these attributes go on the metrics?
Because most of them are unbounded or close to it, and a metric label
that's unbounded turns every distinct value into its own permanent time
series — a well-known way to blow up a metrics backend's cardinality.
mcp.failure.fingerprint and mcp.failure.signature are effectively
unbounded (a new bug is a new fingerprint, forever); mcp.failure.error_class
is medium-cardinality and still risky. Only category (8 values) and
origin (3 values) are small, closed enums — 24 combinations, maximum —
so they're the only two safe to attach to a counter or histogram label.
opentel-mcp enforces this structurally, not by convention. src/metrics.js
can only reach a fingerprint-derived value through one array:
// src/fingerprint/attributes.js
export const METRIC_SAFE_ATTRIBUTES = Object.freeze([
ATTRIBUTE_KEYS.CATEGORY, // 'mcp.failure.category'
ATTRIBUTE_KEYS.ORIGIN, // 'mcp.failure.origin'
]);There's no code path today that lets fingerprint, signature, or
error_class reach a metric label — not a lint rule or a code-review
convention, an actual missing code path. See
Metrics for where mcp.failure.category shows up on
the mcp.tool.errors and mcp.tool.silent_failures counters.
What happens if fingerprinting itself fails?
It falls back to a fixed result rather than throwing or producing a partial, inconsistent fingerprint:
{
"fingerprint": "0000000000000000",
"signature": "unfingerprintable",
"category": "unknown",
"origin": "<whatever origin was passed in>"
}computeFingerprint() runs its entire pipeline inside one top-level
try/catch specifically so this can't surface as an exception — it
executes inline in the tool-call instrumentation path, so throwing here
would mean a bug in fingerprinting could break the tool call it's
trying to observe. This mirrors the same fail-safe design
isToolResultError() and the rest of opentel-mcp's tracing follow — see
Silent Failures.
Can I customize the classifiers or stack depth?
Yes, through computeFingerprint()'s third argument:
computeFingerprint(err, ctx, {
classifiers?: readonly Classifier[]; // prepended to the 7 built-in ones
stackFrames?: number; // default: 5
});Custom classifiers run before the 7 built-in ones, so they get first refusal on an error. Full type signatures are on the API Reference page.
computeFingerprint()'s classifiers and stackFrames options aren't
currently wired through instrumentMcpServer()'s own options — using
them today means calling computeFingerprint() directly rather than
configuring it through instrumentMcpServer(). Still not wired
through as of v0.10.0 — instrumentMcpServer() gained costTracking
(v0.5.0), thrashDetection (v0.6.0), and schemaDrift (v0.8.0) options
in the releases since, but not this one. No target version has been set
for wiring classifiers/stackFrames through
instrumentMcpServer()'s options.
Where do I go from here?
- Silent Failures — the
isError: truedetection everytool_error-origin fingerprint starts from. - Metrics — how
mcp.failure.categoryshows up on themcp.tool.errorsandmcp.tool.silent_failurescounters. - API Reference — full signatures for
computeFingerprint,toSpanAttributes, and the fingerprinting types.