Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Overview

OpenTelemetry for code-mode MCP servers.

It shows you what one of these servers actually did, in the observability stack you already run. If that already makes sense, jump to Install. If not, here’s the whole problem in a page.

The problem

The Model Context Protocol lets an AI agent call tools on your server. Normally it calls one tool at a time. Search for companies. Then enrich this person. One request each.

Code mode is different. Instead of calling one tool, the agent writes a small program and sends you that. You run it in a sandbox, and while it runs, the program calls your tools itself:

const companies = await callTool("inventory_search", { q: "blue widget" });
for (const c of companies.rows) {
  await callTool("item_fetch", { id: c.id });
}

It’s much faster and much cheaper than sending every call back through the model. It’s also much harder to see into.

From outside your server, the whole run is one tool call. One request in, one result out. Which tools the program called, in what order, what it passed, what came back, how long each took, which one broke: all of that happened inside, and none of it got recorded.

agent ──"run this program"──▶ your server ──▶ sandbox ─┐
                                   ▲                   │ callTool("inventory_search", …)
                                   └───────────────────┘ callTool("item_fetch",  …)
                                                         callTool("item_fetch",  …)
      ◀────"here is a result"────

So when a customer says “it gave me the wrong answer”, you’ve got the program and the answer and nothing in between.

The fix

You add two wrappers to your server. From those you get all three OpenTelemetry signals:

  • Traces. One span for the program, one per call it made, nested underneath, with durations and outcomes. Any trace viewer draws it as a waterfall you can read.
  • Metrics. Duration histograms for runs and for calls, so you can ask questions across many runs rather than one.
  • Logs. A record when a run starts, which is the only thing that shows work in flight, because a span doesn’t appear until it finishes.

Each one goes wherever that signal already goes in your setup, and each costs nothing if you don’t run it. No OpenTelemetry at all? You can send everything to your logger instead.

Versus logging

Plenty of teams do, and it works. Three things are hard to get right that way.

Knowing what you can trust. The program is written by an AI, and it can say whatever it likes. If any part of your telemetry comes from what the program printed or threw, then the program is picking what your dashboard shows. mocon tags every value as something your server saw or something the program said, so you can tell them apart. Nothing in OpenTelemetry does this.

Knowing what missing data means. A run with no calls recorded means one of two opposite things. Either the program made no calls, or your server can’t see the calls it made. You declare which, once, and every span carries the answer.

A fixed vocabulary. A run ends one of four ways and a call ends one of three, with the same names on every server that follows this. So a dashboard or an alert or a script you write against those names works on any of them.

Next

Status

Version 0.1.0. Everything here can still change. The attributes mocon defines are its own, and the gen_ai.* and mcp.* ones it reuses are still in development upstream.

Install

npm install @tanvincible/mocon @opentelemetry/api

@opentelemetry/api is a peer dependency, so you pick the version and mocon uses whichever one your application already has.

Python isn’t published yet, so it comes from a clone:

git clone https://github.com/tanvincible/mocon
pip install ./mocon/packages/python

The distribution will be pymocon, because mocon on PyPI is an unrelated project. The import is mocon either way.

Destination

mocon emits through the OpenTelemetry API and never the SDK. That’s on purpose. It means your app decides where telemetry goes, and mocon has no opinion and no config of its own.

It also means nothing comes out until your app registers a tracer provider. If there isn’t one, the OpenTelemetry API quietly does nothing. No spans, no error, no warning, exit code zero. This is the number one reason an integration looks like it isn’t working.

Already running OpenTelemetry? You’re done, go to your first trace.

If you’re not, you’ve got two options.

Set it up. You’ll need an SDK, an exporter, and somewhere for spans to land: Tempo, Jaeger, Honeycomb, Datadog, whatever. That’s real infrastructure work, so decide on it for its own reasons, not because a library asked you to.

Or skip it. If your telemetry today is structured logs, send mocon’s output to the logger you already have. One line, no new infrastructure. See Logging. You can switch to real tracing later without touching your server code.

Requirements

Node 20 or newer. TypeScript types are included, and plain JavaScript works fine.

Quick start

A working integration, start to finish. About ten minutes.

1. Hooks

mocon needs two hooks.

The handler that runs a submitted program. Usually whatever sits behind your execute tool, just before it hands the program to the sandbox.

The function you give the sandbox so it can call your tools. Whatever you inject as callTool or similar. Take the outermost one, the thing the sandbox actually holds.

2. Instance

import { codeMode } from "@tanvincible/mocon";

export const observed = codeMode({
  capabilities: {
    observes_crossings: "some",
    unmediated_egress: true,
    crossing_edge: "invocation",
    attested: [],
  },
});

Those four values say what your server can see. These are the cautious defaults and they’re a safe place to start. Declaring covers how to sharpen them once you’ve checked.

3. Wrappers

return observed.execution.run(
  { program: source, tool: "execute" },
  async (execution) => {
    const callTool = execution.instrument(bridge.callTool);
    return runInSandbox(source, { callTool });
  },
);

That’s both wrappers. The outer one covers the run, and instrument covers every call the program makes through that function.

4. Provider

mocon emits through the OpenTelemetry API and nothing else, so until something registers a provider your spans go to a no-op and you see nothing. That is the usual reason a first run looks silent.

If you already register one somewhere, you are done, skip this. If you don’t, the SDK is a separate install:

npm install @opentelemetry/sdk-node

Then in your real entrypoint, before anything else loads:

import { NodeSDK } from "@opentelemetry/sdk-node";
new NodeSDK({ /* your exporter */ }).start();

Or skip the SDK entirely and write to your logger, which needs no extra install.

5. Run it

Send a program that makes a couple of calls, including one that fails. You should get one execute_code span with two or three execute_tool spans under it.

Output shows exactly what’s on them.

Before shipping

Declare honestly. The defaults above claim almost nothing. Sharpening them is what makes the data worth trusting, and getting it wrong is the one mistake that quietly ruins everything else.

Check how your bridge reports failure. If your callTool returns { ok: false } instead of throwing, mocon will record every failure as a success until you tell it otherwise. One option fixes it, see Wrappers.

Output

One run, two calls, second one refused. Here’s everything mocon produces for it.

Shape

execute_code execute                 303 ms   completed
├── execute_tool inventory_search   127 ms   output
└── execute_tool order_ship          17 ms   error     refused

One span for the program. One span per call it made, nested underneath, in the order they started.

Run span

name                              execute_code execute
kind                              server
duration                          303 ms

code_mode.execution.disposition   completed
code_mode.execution.id            exec_7f3a
code_mode.program.hash            sha256:6719bd29…
code_mode.program.language        javascript

code_mode.observes_crossings      all
code_mode.unmediated_egress       false
code_mode.crossing_edge           invocation
code_mode.attested                ["crossing.target","crossing.input"]

disposition is how the run ended. Read this, not the span’s status. It’s one of completed, failed, terminated (you stopped waiting) or abandoned (you closed the record without finding out). Span status only has three values so it can’t hold all four, which means a run you gave up on looks the same as a clean one in most default dashboards.

execution.id is your own id for the run, the one in your logs. Every span of the run carries it, so one query takes you from a log line to the whole trace.

The last four are the declaration, which is what makes “no calls recorded” mean anything.

A success

name                              execute_tool inventory_search
kind                              client
duration                          127 ms

gen_ai.tool.name                  inventory_search
code_mode.crossing.outcome        output
code_mode.crossing.seq            1
code_mode.crossing.dispatched     true
code_mode.execution.id            exec_7f3a

outcome is output, error or abandoned. Same deal as disposition, read this and not the status, because abandoned and output both look like “unset” to a trace viewer.

dispatched says the call really went out. If your own server refused it, set this false, or whoever’s debugging will go looking in the wrong system.

A failure

name                              execute_tool order_ship
duration                          17 ms
status                            error

code_mode.crossing.outcome        error
error.type                        refused
code_mode.crossing.dispatched     false
code_mode.error.message           "over the call cap for this run"

Provenance

If your server didn’t see a value itself, mocon says so, right next to it:

gen_ai.tool.name                            order_ship
code_mode.provenance.gen_ai.tool.name       P

P means the program said this. No label means your server saw it. There’s also T, which means a target reported it.

This matters because the program is written by an AI. If you build call records out of what the program printed, a program can put a call in your trace that never happened. You can’t tell from the span, so mocon tags it. Provenance has the full story.

Payloads

Program text, call arguments and results are off by default, because they’re AI-written code and customer data. Turn them on when you want them:

codeMode({ capabilities, capture: { values: true } });

Then you also get the arguments and results, cut at a size cap, with a note recording the original size and hash of anything that got shortened. See Payloads.

Signals

The same two wrappers produce all three OpenTelemetry signals. Each one is on by default and costs a function call if your app hasn’t configured that signal, so you get whatever you already run.

Traces

One span per program, one per call it made, nested. The shape of a single run.

Covered in Output.

Metrics

Two histograms, both in seconds, for questions across many runs.

InstrumentKeyed on
code_mode.execution.durationdisposition, and error type when there is one
code_mode.crossing.durationtool name, outcome and error type

The second one is conditional, and this is the interesting part. Those dimensions are only added when you attested crossing.target. If you didn’t, the tool name is whatever the program said it called, and a metric has nowhere to record that doubt, so it would turn a claim into a fact that nothing downstream could question. So the dimensions get dropped and you get an undimensioned duration distribution instead. Less useful, not a lie.

Calls that ended abandoned are not recorded at all. Their duration is zero by construction, so counting them would put a fiction in the distribution.

You don’t need a collector for these. They come straight from your app through whatever metrics exporter you already have. The collector is still worth running if something else in your pipeline derives metrics from span names, because that you can’t control from here.

Logs

One record when a run starts, one when it ends.

{
  "event.name": "code_mode.execution.started",
  "body": "a program dispatch started",
  "trace_id": "7269fe4c…",
  "span_id": "a1fa92d3…",
  "code_mode.execution.id": "run-1"
}

The starting record is the point. A span only exports when it ends, so a run that’s still going, or one that hung, isn’t in your trace at all. It looks exactly like a run that never happened. The log record is the only thing in the whole model that says a run is in flight right now.

Both records carry the trace and span id, so they’re a view of the trace rather than a second source of truth. Join on those ids and you’re back in the waterfall.

This needs @opentelemetry/api-logs, which is an optional peer dependency. If it isn’t installed, log records are silently skipped and everything else works.

Switching off

codeMode({
  capabilities,
  signals: { metrics: false, logs: false },
});

Traces always emit. The other two are on unless you say otherwise.

No OpenTelemetry

If you run none of this, use your logger instead. You lose metrics and the waterfall, and keep the vocabulary.

Wrappers

The run

observed.execution.run({ program: source, tool: "execute" }, async (execution) => {
  // your existing handler body
});

Start it before your first rejection. A run starts when you first see the submission, including ones you then refuse for a bad key, a failed lint, or being at capacity. If you start the span after those checks, every refused run is invisible, and an outage that rejects everything looks exactly like no traffic.

For a refusal, end it on purpose:

execution.fail(new Error("unknown tool in script"), { errorType: "validation" });

If your handler returns a failure object instead of throwing, say so, or every failed run gets recorded as a success:

observed.execution.run(
  {
    program: source,
    end: (value) => (value.ok ? undefined : { disposition: "failed", errorType: "runtime" }),
  },
  body,
);

Return undefined from end to mean “just use the default”.

The bridge

const callTool = execution.instrument(bridge.callTool);

Wrap the outermost function the program can reach. If your sandbox gets a function that calls another one that then dispatches, wrap the one the sandbox holds. Anything above your wrapper that can answer the program produces no span at all.

It has to be your code, outside the sandbox. A function the program can reach, swap out or watch is just another thing the program controls. If yours is reachable from inside, your telemetry says whatever the program wants.

No function to wrap? If your sandbox runs somewhere else and reports back, see No bridge. You record the calls yourself and get the same spans.

Envelopes

Lots of bridges return { ok: false, error } instead of throwing. mocon reads a normal return as success, so on a bridge like that every failure gets quietly recorded as working. One option fixes it. Return only what you want changed: everything you leave out is filled in from what the bridge actually answered, so the envelope still lands on the span as the reason.

const callTool = execution.instrument(bridge.callTool, {
  end: (answer) =>
    !answer.threw && !answer.value.ok
      ? { outcome: "error", errorType: "capability_error", dispatched: true }
      : undefined,
});

Extra arguments

By default the first argument is the target and everything after it is the input. So a bridge shaped callTool(name, params, { signal, deadline }) ends up recording your own abort signal and deadline as the program’s arguments. Tell it what the input really is:

execution.instrument(bridge.callTool, { input: (_name, params) => params });

Local calls

Span kind defaults to client, which says you forwarded the call somewhere remote. For a tool your own process serves, say so:

execution.crossing.start({ target: "cache_get", kind: "local" });

By hand

instrument covers the normal case. When you need more control, open and close a call yourself:

const crossing = execution.crossing.start({ target: "inventory_search", input: params });
try {
  const result = await dispatch(params);
  crossing.output(result, { dispatched: true });
} catch (e) {
  crossing.error(e, { errorType: "capability_error", dispatched: true });
}

A call you never close gets closed for you as abandoned when the run ends, so nothing dangles.

No bridge

instrument wraps a function, which only helps if your host has one to wrap. Plenty don’t.

Maybe your sandbox runs in another process and calls back over HTTP. Maybe it drops requests on a queue and something else picks them up. Maybe you only find out what it did by reading a log after it finishes. In all of those there’s no function sitting on the boundary, so you record the calls yourself.

You get the same spans either way. instrument is a convenience built on top of what’s below.

Recording

Three calls. Start one when you learn a call began, end it when you learn how it went.

const crossing = execution.crossing.start({ target: "orders.list", input: params });

crossing.output(result);  // it came back
crossing.error(failure);  // it didn't

target is the only thing required. Everything else is optional and gets filled in with what you know.

Late

Calls don’t have to end in the order they started, and they don’t have to end at all.

const a = execution.crossing.start({ target: "orders.list" });
const b = execution.crossing.start({ target: "inventory.check" });

b.error(new Error("upstream 503"));   // second one settles first
a.output({ rows: 2 });

execution.crossing.start({ target: "slow_thing" });  // never settles
execution.complete();
execute_tool inventory.check   error
execute_tool orders.list       output
execute_tool slow_thing        abandoned
execute_code                   completed

Anything still open when the execution ends gets closed for you and marked abandoned. A sandbox that goes quiet leaves a record saying so rather than a hole you have to notice.

Order

Timestamps often won’t order these for you. Calls under a millisecond tie, and a sandbox on another machine has a clock you don’t control. If you know the order the program asked in, say it:

execution.crossing.start({ target: "orders.list", seq: 2 });

Only pass seq if it’s real. A wrong order is worse than no order, so anything that isn’t a positive integer is refused and the count mocon kept itself is used instead.

Declaring

Watching from a distance usually means seeing less, and the declaration is where you say so. Three settings matter here.

crossing_edge. Use "dispatch" if what you see is the request you sent toward the tool, rather than the call the program asked for. Anything observing at the network layer is dispatch. A retry you did on the program’s behalf is one call to the program and several dispatches, and this is what tells a reader which one they’re looking at.

observes_crossings. Use "some" unless you’re certain you see every call. If calls reach you through a queue you might drop from, or a log you might read late, you see some.

unmediated_egress. Set it true if the program has any path out that doesn’t come through you. This is the one that stops a reader concluding “no calls recorded, so it called nothing”, which is the wrong conclusion to let someone reach.

codeMode({
  capabilities: {
    observes_crossings: "some",
    unmediated_egress: true,
    crossing_edge: "dispatch",
    attested: [],
  },
});

Declaring goes through all of it, including how to sharpen these once you’ve checked what you actually see.

Reconnecting

If your host learns about an execution in one place and its calls in another, you don’t need to hold the handle between them. Give the execution an id you already have:

const execution = observed.execution.start({ program: source, id: runId });

Every crossing carries that id, so a query finds them by it without needing the parent span. That’s what lets a worker process record a call for a run that a different process started.

Declaring

Four values, set once, that ride on every span. They tell whoever’s reading how much of the picture they’re actually looking at.

codeMode({
  capabilities: {
    observes_crossings: "all",
    unmediated_egress: false,
    crossing_edge: "invocation",
    attested: ["crossing.target", "crossing.input", "crossing.output"],
  },
});

Why

A run whose trace shows no calls means one of two opposite things. Either the program made no calls, or your server can’t see the ones it made. Nothing else in the trace tells them apart. This does, and every other claim depends on it being honest.

observes_crossings

all, some, or none.

all means nothing can answer the program before your wrapper does. Not “my wrapper sees every call that reaches it”. Before you claim it, go look for code that answers the program itself:

  • a cap on calls per run
  • a deadline or time budget guard
  • a rate limiter
  • a cache that returns without dispatching
  • a permission check that refuses before dispatch

If any of those can return to the program without going through the function you wrapped, then some calls make no span, and all is false. Use some.

How to check, in five minutes. Instrument it, write a program that deliberately hits every refusal path you’ve got, and count the spans against the calls. If they don’t match, you’re some.

unmediated_egress

true if the program has any way out that you don’t see: raw network, subprocesses, an isolate that can be escaped. It stops someone concluding “three spans, so three external calls”.

Not sure your sandbox is airtight? true is the honest answer.

crossing_edge

invocation if a span describes what the program asked for. dispatch if it describes what you actually sent after retries and rewrites. Most integrations wrap the bridge the program calls, so that’s invocation.

attested

By default everything is treated as a program claim, which is the safe reading. This list is how you upgrade specific things to “my server saw this”:

EntryWhat it upgrades
crossing.targetthe tool name, its order and its outcome
crossing.inputthe call arguments
crossing.outputthe result, to “a target reported it”
crossing.errorthe error class and message, to “a target reported it”
execution.error.classthe run’s error type
host_attributesyour own attributes, listed separately

Only attest something if it’s true for every span you emit. There’s no per-call opt-out.

Don’t attest anything you work out from what the program wrote. If your error class comes partly from matching a thrown value’s name or message, the program can pick it. If you’ve got both an observed path and a parsed path for the same field, don’t attest that field.

The rule

Declare the weakest thing that’s true for every run. Saying nothing reads as none, nothing attested, egress unknown, and that’s safe. Forgetting to claim something costs you a bit of detail. Claiming something that isn’t true quietly corrupts every conclusion anyone draws from your traces.

Provenance

The program running in your sandbox was written by an AI. It can print anything, throw anything and return anything.

That matters because of where telemetry comes from. If you record a call because the program logged one, then a program that logs a call it never made just put a fiction in your trace. If you classify errors by matching the message, then a program throwing new Error("timeout") picked what your dashboard says.

This isn’t about hostile programs. It’s just what happens when you build telemetry out of something the subject controls.

And more and more, the thing reading these traces is another model. A model reading “the program deleted contact 42” needs to know whether you watched that happen or the program said it did.

No general observability tool records this. A span attribute is a key and a value. There’s nowhere to put where the value came from. That’s the gap this fills.

Three classes

Your server saw it. Determined somewhere the program can’t write: your clock, your id generator, an exit status, a call boundary you control.

The program said it. Written by the program, or worked out by you from something it wrote: the program text, its output, its thrown errors, its return values.

A target reported it. Passed through unchanged from whatever the call reached, or produced by your own handling of that call, like a refusal. The program didn’t shape it.

Labels

Any value that isn’t something your server saw gets a label right next to it:

gen_ai.tool.name                            inventory_search
code_mode.provenance.gen_ai.tool.name       P

P for the program said it, T for a target reported it, and no label for your server saw it.

The default runs the safe way round. Anything you haven’t declared you observed gets P. So forgetting to declare something costs you a bit of detail, and it can’t accidentally turn a guess into a fact.

A trap

T doesn’t mean the target saw the call. A refusal your own server produced is T, because the program didn’t shape it. Someone reading a T error will naturally go digging in the target’s logs for a request that never left your process.

code_mode.crossing.dispatched is what separates them. Set it, and “their API broke” versus “we never called them” is one field instead of an afternoon.

Not proof

Saying you observed something makes the claim visible and makes it yours. It doesn’t make it true. Nothing in a trace can tell a server reading its own call boundary apart from a server copying a value out of the program’s return and attesting it anyway.

No format can catch that. What a format can do is put a name on the claim, so if it’s wrong, it’s wrong in public.

Unlabelled

A span’s name and its status description have no attribute key, so nothing can sit beside them.

The name matters because span-metrics tools, service maps and name-keyed alerts all read it. On a server that doesn’t attest its targets, they’re reading a program’s claim as fact. The collector stops the metrics mocon defines from doing that. It can’t stop tooling somebody else set up.

The status description is handled by never putting anything unlabellable in it. It carries a fixed-vocabulary value, never the program’s words. The message goes in an attribute, where it can be labelled.

Custom attributes

Your server knows things mocon doesn’t. Credits spent, a sandbox id, an attempt count, a cache hit, a tenant. Put them on the spans.

execution.crossing.start({
  target: "inventory_search",
  attributes: { "com.acme.credits_used": 5, "com.acme.cache_hit": false },
});

Use your own namespace, from your domain or product name. Keys under code_mode., gen_ai., mcp. and otel. are reserved and get dropped rather than written.

Trust

By default everything you add gets labelled unverified, because mocon has no idea where your value came from. Two lists sort that out:

capabilities: {
  attested: [...otherEntries, "host_attributes"],
  attested_attributes: ["com.acme.sandbox_id"],    // you measured these
  relayed_attributes: ["com.acme.credits_used"],   // a target reported these
}

attested_attributes is for things you worked out yourself, somewhere the program can’t reach. Your own meter, your own clock, your own sandbox id.

relayed_attributes is for numbers you copied out of a target’s response. A credit count your API returned isn’t something you measured. Attesting it would be a lie, and leaving it off makes it a program claim, which bars it from becoming a metric. Listing it as relayed is the honest option, and the one that gets you a billing number you can actually defend.

A key can’t be in both lists, and listing keys at all needs host_attributes in attested.

Meaning

Nothing outside your server knows what com.acme.credits_used is. Tell it:

capabilities: {
  declared: {
    "com.acme.credits_used": { agg: "sum", unit: "{credit}", card: "low", name: "Credits" },
    "com.acme.tenant_id":    { agg: "none", card: "high" },
  },
}
FieldWhat it says
aggsum if the values add up, last if only the newest matters, none if adding them is meaningless
unitUCUM if there is one (ms, By, s), otherwise a braced annotation ({credit})
cardlow if it’s safe to group by, high for per-user or per-run values
namewhat to call it in a legend

Summable is a different question from believable, and a reader needs both answers. The declaration says the number adds up. The provenance label says whose number it is. A value declared summable that carries a P label still shouldn’t become a metric, because a metric has nowhere to carry the doubt.

Readers

Worth being straight, because self-describing data is easy to oversell.

No general-purpose backend reads it. Grafana isn’t going to learn what your credit meter is because a span told it.

Three readers get something out of it. A model reading the trace, which can act on it with no vendor support at all. A dashboard written against mocon, which can then render your fields without being rebuilt for your server. And a collector deriving metrics, which can follow the provenance rule automatically instead of trusting each server.

So it doesn’t make the ecosystem understand you. It makes it possible for something to.

Payloads

Program text, call arguments, results and error bodies are off by default. They’re AI-written code and customer data, so you opt in rather than out.

codeMode({
  capabilities,
  capture: { values: true },
});

Captured

AttributeWhat it is
code_mode.program.textthe program that was submitted
gen_ai.tool.call.argumentswhat the program passed to a call
gen_ai.tool.call.resultwhat came back
code_mode.error.messagethe human-readable reason a thing failed
code_mode.error.bodythe raw error object
code_mode.output.<channel>stdout, stderr, whatever else you capture

The program’s hash is always written, capture on or off. That’s how you tell two runs of the same program apart, and it’s what’s left when you withhold the text itself.

Size caps

Big values get cut:

capture: {
  values: true,
  cap: 8192,          // bytes written per value
  programCap: 32768,  // bytes written for the program
  measure: 1048576,   // bytes read to work out the real size and hash
}

When something gets cut, a note records what happened:

code_mode.capture  {"gen_ai.tool.call.result":{"truncated":true,"bytes":62240,"hash":"sha256:…"}}

So you still know the real size and can match it against the full value elsewhere. OpenTelemetry has no way to say “this value was shortened”, which is why the note exists.

Set your cap below whatever limits your SDK, collector and backend have. Better to cut it yourself and say so than to have something downstream cut it silently.

Unserializable

A value with a cycle in it, a getter that throws, a toJSON that blows up: these all get recorded as redacted rather than crashing anything.

code_mode.capture  {"gen_ai.tool.call.result":{"redacted":true}}

That matters more here than in most libraries, because the values come from AI-written code. Whatever the program returns, capturing it can cost you the value and never the call.

Two more things land here. A NaN or an infinity anywhere in a value redacts the whole value, since JSON has no way to write either and putting null there would turn a reading into a reading of nothing. And a single value far past your cap is refused rather than read, because serializing something enormous is work a program can ask for without limit.

A redacted note carries no bytes and no hash. Both describe an original the host never managed to serialize, so there is no honest number to report.

Redacted

Absent means you never captured that thing. It says nothing.

Redacted means you had it and removed it on purpose. That’s a real signal, so they’re different on the wire.

To keep a program’s hash without its text, the shape is: no code_mode.program.text, a present code_mode.program.hash, and a capture note saying redacted.

Privacy

Everything here is program-determined, so it all carries a P provenance label. If you’re sending traces somewhere you don’t fully control, values: false gives you the whole structure of a run, durations, outcomes, ordering, failures, with none of the content.

Logging

Standing up a collector and a trace backend is a real decision. If your telemetry today is structured logs, that’s a lot more work than the two wrappers. You don’t have to do it.

import { codeMode, logTracer } from "@tanvincible/mocon";

const observed = codeMode({
  capabilities: { /* same as before */ },
  tracer: logTracer((record) => logger.info(record)),
});

That’s the only line that changes. No SDK, no exporter, no collector, no backend.

Records

Every finished span becomes one flat record handed to your logger:

{
  "name": "execute_tool inventory_search",
  "kind": "client",
  "trace_id": "f3d8f44c4f3d6a83bd2518356040dd1a",
  "span_id": "33924c3b939b1dcc",
  "parent_span_id": "eceff414b8114181",
  "start": "2026-09-20T10:14:02.118Z",
  "duration_ms": 17.68,
  "code_mode.execution.id": "exec_7f3a",
  "code_mode.crossing.outcome": "output",
  "gen_ai.tool.name": "inventory_search",
  "gen_ai.tool.call.arguments": { "q": "widget" },
  "code_mode.provenance.gen_ai.tool.call.result": "P"
}

The whole attribute set, the provenance labels, the ids, a duration and a status. Payloads come back as real values rather than JSON strings, because a log record can hold an object where a span attribute can’t. Group by code_mode.execution.id and you’ve got the whole run, in the pipeline you already query.

Trade-offs

The things a trace store is actually for. A rendered waterfall, and metrics off spans without aggregating log lines yourself.

Reversible

The ability to change your mind. Switching to a real trace pipeline later means passing a different tracer and touching nothing else.

Flat

Records are flat, one per span, rather than nesting calls inside their run.

Nesting means holding children until the parent closes, and a call the program makes a tick later then never gets written at all. Flat records carry parent_span_id, so you rebuild the tree by grouping instead of trusting the writer to buffer correctly.

Options

logTracer({
  write: (record) => logger.info(record),
  raw: true,   // leave payloads as JSON strings instead of decoding them
});

If your logger throws, the record is dropped and the call carries on. An observability problem should never break the thing it’s watching.

Mistakes

Five things that go wrong, roughly in order of how much damage they do.

Over-declaring

observes_crossings: "all" means nothing can answer the program before your wrapper. If your sandbox refuses calls over a cap, or a deadline guard answers early, or a cache returns without dispatching, those calls make no span and "all" is a lie.

It’s the worst one because everything else rests on it, and you can’t spot it afterwards from the data. Four calls, two spans, and a declaration saying two was all of them looks completely normal.

Fix: write a program that hits every refusal path you have, and count spans against calls. If they don’t match, you’re "some". Declaring.

No provider

The OpenTelemetry API does nothing when no provider is registered. No spans, no error, no warning, exit code zero. Registering one in a test or a demo script doesn’t count.

Fix: grep your own src/ for NodeSDK or TracerProvider and make sure you find something outside a test. Or use your logger, which needs no provider at all.

Silent failures

If callTool returns { ok: false } instead of throwing, the default reads that as success. Every failed call gets recorded as working, and the trace looks healthy while your users don’t.

Fix: the end option on instrument. See Wrappers.

Late start

Runs you refuse for a bad key, a failed lint or being at capacity never produce a span at all. The failure modes you most want to see are the ones that vanish, and they vanish in a way that looks like nobody called you.

Fix: start the run span first, then execution.fail(cause, { errorType: "validation" }).

Leaked internals

The default treats every argument after the first as the program’s input. A bridge shaped callTool(name, params, { signal, deadline }) therefore records your abort signal and deadline as things the program passed, and attesting crossing.input publishes that as fact.

Fix: input: (_name, params) => params.

Smaller

Leaving kind unset is a choice. It defaults to client, which says you forwarded the call somewhere remote. Pass kind: "local" for tools your own process serves.

No context manager means neighbouring instrumentation floats. mocon puts the run span in the active context so other instrumentation nests under it, but that only works if your app registered a context manager. NodeSDK does. A hand-assembled provider doesn’t. mocon’s own spans are fine either way, which is exactly why it’s easy to miss: your trace looks perfect and everything else drifts off.

Rollout

The code change is the small part. If you’re replacing existing observability, the order you do things in decides whether the day you merge is better or worse than the day before.

Order

1. Decide if you want a trace pipeline. Already running OpenTelemetry? Most of the cost is already paid. If your telemetry goes to logs, adopting this means running a collector and a trace store, and that’s a decision to make on its own merits. If the answer is no, use your logger and you’re done. You lose the waterfall and keep everything else.

2. Stand up the destination first. Merge the collector into your collector, point it at your backend, and check data arrives with nothing instrumented yet.

3. Import the dashboard. dashboards/code-mode.json, repointed at your datasources. Make sure it renders empty rather than broken.

4. Then the code. Two wrappers, a provider in your real entrypoint, a cautious declaration.

5. Turn it on in staging. Run a program that makes several calls, including one that fails and one your server refuses. Check the spans arrive and the dashboard fills.

6. Retire the old thing last. Deleting a working log line in the same change that ships its replacement switched off leaves you worse off than before you started. Wait until the new data is confirmed flowing.

Sampling

Every run emits one span plus one per call, so a program making a hundred calls emits a hundred and one. Configure a sampler before this becomes a bill.

Use a parent-based sampler so a run and its calls are kept or dropped together. With a head sampler that decides per span you get runs with half their calls missing, and a missing call is indistinguishable from a call that never happened.

Cost

mocon adds roughly five microseconds per span on top of what the OpenTelemetry SDK costs, and about eight with payload capture on. A run executes a whole program and a call is usually a network request, so this sits far below the work it’s describing. It holds nothing between runs.

Run npm run bench in the repo if you want your own numbers.

Reverting

Take the wrappers out, or leave them and don’t register a provider. Everything becomes a no-op with no other changes. If you used the logger, swap the tracer back.

Collector

Optional. The two histograms now come from your app directly, so you do not need a collector to get metrics at all. This is for one narrower job your server genuinely cannot do for itself. It stops a program’s claim from turning into a metric that looks like a measured fact.

If you don’t attest crossing.target, the tool name on a call span is whatever the program said. A span-metrics connector doesn’t read provenance, so left alone it happily produces calls_total{gen_ai.tool.name="order_ship"} from a name the program picked, and a metric has nowhere to carry the doubt.

Your server can’t prevent that, because the connector runs downstream. A collector can, because it sits after every server and before every backend.

Config

On purpose. A custom collector component has to be compiled into a distribution, so everyone adopting it has to rebuild and redeploy their collector first. Everything here is stock opentelemetry-collector-contrib, so it works with the collector you already run.

Pipelines

Two passes over the same spans.

The trace pipeline keeps everything, claims included. A claim belongs in a trace, next to the label saying what it is, where a person reads it in context.

A second traces pipeline feeds the metrics connector and drops the claims first, so nothing unobserved ever gets counted.

Spans with no provenance label are left alone, so telemetry from the rest of your system passes through untouched.

There’s also an off-by-default transform processor that strips program-written payload values, for when you want the shape of a run in your backend but not the content. The capture note survives it, so you still see the size and hash of whatever got removed.

Checking

collector/check.mjs sends two calls through a real collector. One from a server that watched its own boundary, one from a server that didn’t.

in tracesin metrics
inventory_search, observedyesyes
order_ship, claimedyesno

Takes about a minute.

Gotcha

Since v0.104 the collector binds OTLP to localhost by default, so a collector in a container with the stock config receives nothing at all and says nothing about it. The config here sets explicit endpoints.

Dashboard

dashboards/code-mode.json is an example, not the product. It happens to be Grafana because that is what it was built against. Use whatever you already run.

Why example

Everything here is ordinary OpenTelemetry. The two histograms come out named, united and described, so they show up correctly in any metric browser without anyone teaching it anything. The spans are spans. The log records are log records. Datadog, Honeycomb, Grafana, Elastic and the rest all handle them the same way they handle everything else you send.

So there is nothing to build before you can look at this. There is only a choice about what you want on one screen, and that is yours rather than ours.

Views

If you are building your own, these are the views that show something the raw trace does not.

Can you believe there were no calls? Group runs by code_mode.observes_crossings and code_mode.unmediated_egress. Watching everything with no unmediated egress is the only combination where a run showing no calls really made none.

Runs by disposition. code_mode.execution.duration grouped by code_mode.execution.disposition. Span status has three values where this has four, so a run you gave up on looks identical to a clean one everywhere else.

Calls by outcome. Same idea on code_mode.crossing.duration and code_mode.crossing.outcome. abandoned and output are both “unset” to a trace viewer.

Abandoned calls. Count of crossings with that outcome. If your targets spend money or change state, that is how many things may or may not have happened.

Calls the program claimed. Search traces for spans carrying code_mode.provenance.gen_ai.tool.name. By design these reach no metric, so a trace view is the only place they appear.

Runs in flight. Log records with event.name = code_mode.execution.started that have no matching ended. Nothing else in the model can show work still running.

Ours

Import dashboards/code-mode.json, repoint its two datasource uids, and you get the six views above in Grafana. The PromQL is tested. The four trace panels are not, so give them a look.

It needs the metrics, which come from your app directly, and a Tempo-compatible trace store for the trace panels.

API

codeMode(options)

Makes an instance. Do this once, at startup.

codeMode({
  capabilities,      // required, see below
  tracer,            // optional, defaults to the global OpenTelemetry tracer
  capture,           // optional, payload settings
});

Bad capabilities throw here, at startup, rather than later on a request.

capabilities

FieldType
observes_crossings"all" | "some" | "none"required
unmediated_egressbooleanrequired
crossing_edge"invocation" | "dispatch"required unless observes_crossings is "none"
attestedstring[]what your server observed, default []
attested_attributesstring[]your keys you measured, needs host_attributes
relayed_attributesstring[]your keys a target reported, needs host_attributes
declaredobjectwhat your keys mean

capture

FieldDefault
valuesfalsewrite program text, arguments and results
cap8192bytes written per value
programCap32768bytes written for the program
measure1048576bytes read to compute the real size and hash

execution.run(options, body)

Starts a run, calls body, closes the run. Returns whatever body returns, and follows a promise if it returns one. A throw is recorded as failed and rethrown unchanged.

observed.execution.run({ program, tool: "execute" }, (execution) => { … });
Option
programthe submitted text, required
toolthe name of your code-mode tool
idyour own run id; one is generated if you don’t pass one
languagea hint like "javascript", leave it out rather than guess
kind"server" (default) or "local"
parentthe caller’s context from your propagator, never from the sandbox
sessionId, conversationId, toolCallIdcorrelation ids
attributesyour own attributes
endread a failure envelope, see the wrappers

execution.start(options)

Same options, but you close it yourself. Use it when your handler shape doesn’t suit a callback.

Execution handle

instrument(fn, options?)wrap a bridge function, one call becomes one span
crossing.start(options)open a call by hand
complete(options?)close the run as completed
fail(cause, options?)close it as failed
end(options)close it with any disposition
spanthe underlying OpenTelemetry span
contextthe run’s context, for bridges served in another task

Closing twice is a no-op, so the first close wins.

instrument(fn, options?)

Returns a wrapped function with the same name, arity and behaviour. It forwards this, rethrows the exact error, and follows a returned promise.

Option
targeta string, or a function of the arguments; defaults to the first argument
inputa function of the arguments; defaults to everything after the target
endturn the bridge’s answer into an outcome
toolType"function", "extension" or "datastore"
attributesyour own attributes

If one of these options throws, you lose that field and not the call. The call still runs and the span is still recorded.

Crossing handle

output(value?, options?)settled with a result
error(cause, options?)settled with an error
end(options)settled with any outcome
spanthe underlying span

Options take dispatched, errorType, message, endTime and attributes. A call you never settle is closed as abandoned when the run ends.

logTracer(write | options)

A tracer that writes flat records to a function instead of exporting spans. See using your logger.

Errors

Bad configuration throws at startup: TypeError for a wrong type, RangeError for a value outside a fixed set.

On the request path, nothing throws. A value that can’t be serialized is recorded as redacted, a broken option costs that field, a logger that throws costs that record. Observability should never break the thing it’s watching.

Attributes

Everything mocon writes. The specification has the normative detail.

Both spans

Attribute
code_mode.observes_crossingsall, some or none
code_mode.unmediated_egresscan the program get out without you seeing
code_mode.crossing_edgeinvocation or dispatch
code_mode.attestedwhat your server observed
code_mode.attested_attributesyour keys you measured
code_mode.relayed_attributesyour keys a target reported
code_mode.execution.idyour own run id, on every span of the run
code_mode.capturewhat got truncated or redacted
code_mode.provenance.<key>P or T for any value your server didn’t observe

Run span

Attribute
gen_ai.operation.namealways execute_code
code_mode.execution.dispositioncompleted, failed, terminated, abandoned
code_mode.program.hashsha256 of the program, always written
code_mode.program.languagea display hint, omit rather than guess
code_mode.program.textthe program, opt-in
code_mode.declaredwhat your own attributes mean
code_mode.output.<channel>stdout, stderr and so on, opt-in
gen_ai.tool.nameyour code-mode tool’s name
gen_ai.tool.call.idthe caller’s id for this dispatch
gen_ai.conversation.idonly if your grouping really is a conversation
mcp.session.idthe MCP session
error.typewhen the run failed

Call span

Attribute
gen_ai.operation.namealways execute_tool
gen_ai.tool.namethe target
code_mode.crossing.outcomeoutput, error, abandoned
code_mode.crossing.dispatcheddid the call actually leave
code_mode.crossing.seqorder it started in, from 1
code_mode.crossing.timingset when a time had to be made up
gen_ai.tool.call.idyour own id for this call
gen_ai.tool.typefunction, extension or datastore
gen_ai.tool.call.argumentsthe input, opt-in
gen_ai.tool.call.resultthe result, opt-in
code_mode.error.messagewhy it failed, opt-in
code_mode.error.bodythe raw error, opt-in
mcp.method.name, mcp.resource.uriif the call went over MCP
error.typewhen it failed

Metrics

InstrumentUnitKeyed on
code_mode.execution.durationscode_mode.execution.disposition, error.type
code_mode.crossing.durationsgen_ai.tool.name, code_mode.crossing.outcome, error.type, only when the target is attested

Log records

Attribute
event.namecode_mode.execution.started or code_mode.execution.ended
trace_id, span_idjoin back to the spans

Plus every attribute the run span carries.

Values

Disposition is completed, failed, terminated, abandoned.

Outcome is output, error, abandoned.

Error types on a run: validation, runtime, timeout, resource_limit, cancelled, approval_rejected, host_failure, _OTHER.

Error types on a call: capability_error, validation, refused, timeout, cancelled, approval_rejected, tool_error, _OTHER.

Both lists of error types are open, so use a more specific low-cardinality name if you have one. The disposition and outcome lists are closed and nothing else is allowed.

Names

NameKind
Runexecute_code {tool}server, or internal in-process
Callexecute_tool {target}client, or internal if you serve it

If your targets are unbounded, like URLs, pass a name to keep the span name low cardinality and leave the full target in gen_ai.tool.name.

Status

Status is a display hint. Read the disposition and outcome attributes instead.

Status
completed, abandoned, outputunset
failed, terminated, errorerror

The description carries error.type, never the program’s words, because it’s the one field nothing can label.

Limits

Things this genuinely can’t do. Worth knowing before you rely on it.

Live runs

A span only exports when it ends. So a run that’s still going, or one that hung, isn’t there at all, and it looks exactly like a run that never happened.

“Which run is stuck right now” is not a question you can answer this way. If you need that, close long-running work as abandoned on a timer in your own server, or emit a log line at start.

This is a real step back from plain logging, which writes things as they happen.

Status

Span status has three values and instrumentation shouldn’t set ok, so in practice you get two. completed and abandoned both look unset. failed and terminated both look like errors.

Every default dashboard reads that field. Yours needs to read code_mode.execution.disposition instead, which is what the dashboard does.

Span names

If you don’t attest crossing.target, the call span’s name is whatever the program said it called. Span-metrics tools, service maps and name-keyed alerts all key on span names, and none of them read provenance.

Collector blocks the metrics mocon defines from doing this. It can’t stop a connector somebody else configured.

Unknown duration

A span always has a start and an end, so it always has a duration. A call you know settled but can’t time becomes a zero-duration span rendering as a tick. There’s an attribute saying so, and no trace viewer reads it.

Attestation

Nothing in a trace can tell apart a server genuinely watching its call boundary from one copying a value out of the program’s return and attesting it anyway. Catching that needs a second observer in the path under its own identity. No format does it.

Fixed sets

You can add a field next to a disposition. You can’t add a value to it, and a reader following the rules treats the fixed value as the real answer.

The case that bites: a run that pauses at the end of one dispatch and picks up in a later one. Each dispatch is its own run, so the paused one reports completed, and one logical run looks like three completed ones.

Outside

Sampling can drop part of a trace, so a missing call might mean sampled rather than never happened. Use a parent-based sampler.

Attribute length limits in the SDK cut values after mocon has already recorded what it did, so something captured whole can arrive shortened with nothing saying so. Set your own cap lower.

Runtime

The TypeScript package needs Node 20 or later. It uses Buffer, node:crypto and node:util, so it won’t run on Workers, Deno, or in a browser. If your sandbox lives on one of those, that package isn’t an option today.

The attributes themselves don’t care. They’re plain OpenTelemetry, the reference lists every one, and the specification says exactly what each means. Emitting them yourself from whatever runtime you’re on gets you the same spans, and a second implementation already does exactly that.

Not standard

code_mode.* is this project’s own namespace and nobody else has agreed to it. The gen_ai.* and mcp.* attributes it reuses are still in development upstream with no compatibility promise, and gen_ai.operation.name = execute_code isn’t an upstream value, so anything filtering on known operation names won’t see these runs at all.

Specification

The normative document, in full. Everything else in this book explains it.

You don’t need this to use mocon. Read it if you’re implementing the specification in another language, writing a consumer, or want the exact rules behind something.

Code-mode specification

Status: Development. Version 0.1.0, 2026-09-19.

Everything here is Development and may change. The gen_ai.* and mcp.* attributes this document reuses are themselves Development, in open-telemetry/semantic-conventions-genai, read at 2026-09-19. Nothing in either namespace is Stable and neither carries a compatibility guarantee. error.type is the one Stable attribute used below, and it belongs to the core semantic-conventions registry, not to gen_ai. A host pins the registry version it built against by setting schema_url on its instrumentation scope once one is published, and until then by setting the scope version to the version of this document.

This project previously specified a JSON Lines record format with its own wire, schema and conformance suite. This document replaces it. The model, the closed vocabularies, the provenance rules and the capability declaration survive that move unchanged; section 14 lists what did not, and Appendix A carries the invariants the whole design rests on.

1. Scope

A code-mode server is one where the agent submits a program instead of one structured tool call. The host runs the program in an environment it controls, and from inside the program the server’s capabilities are reached through a mechanism the host provides. From outside the host that whole run is one opaque tool call: the calls the program made, what it was given, what came back, and whether the host could see any of it are all invisible.

This document says how a host makes that visible in OpenTelemetry. It defines two spans, a declaration of what the host can and cannot observe, a rule for telling a value the host measured from one the program claimed, and a rule for values the host shortened or removed.

It does not define a wire format. OpenTelemetry is the wire format. It does not define a library. It defines what a host emits.

Instrumentation depends on the OpenTelemetry API only, never the SDK. That is OpenTelemetry’s own rule for instrumentation (specification/library-guidelines.md: “Third party libraries and frameworks that add instrumentation to their code will have a dependency only on the API of OpenTelemetry client”), and it is the reason this works: the host emits through the API, and the application owner’s configured exporters receive it. A host that cannot carry an SDK writes OTLP/JSON directly; that encoding is Stable and documented.

Everything in this document is written so that an API-only emitter can produce it. Where that rules a design out, it is ruled out, and section 3 is where it bites hardest.

Namespace. New attributes are under code_mode., lowercase and dot-delimited, snake_case inside each component, per the OpenTelemetry attribute naming rules. gen_ai.* and mcp.* are existing OpenTelemetry namespaces: this document reuses attributes from them and mints nothing inside them, because the naming rules say not to squat an existing convention namespace. otel.* is reserved to the OpenTelemetry specification. The capability attributes are code_mode.observes_crossings and not code_mode.host.*, because host.* in OpenTelemetry already means the machine.

Requirement levels are OpenTelemetry’s: Required, Conditionally Required, Recommended, Opt-In. Opt-In means off by default and turned on by the application owner. It is used here for every attribute that carries program text, call arguments or call results.

The key words MUST, MUST NOT, SHOULD, SHOULD NOT and MAY are to be interpreted as described in RFC 2119.

2. The model

Three terms, and they are the whole model.

Execution. One dispatch of one program by a host. Never the session, the conversation or the container that holds it. A program the host runs again, from a retry, a replay, a resumed checkpoint, or a speculative branch its own substrate took, is a new execution.

Crossing. One invocation, initiated by the program, that crosses from the program to the host-provided surface: a tool call, a binding method, a proxied fetch, a file read the host serves.

Target. The host-defined identifier of what a crossing invoked.

Two closed vocabularies. An execution ends completed, failed, terminated or abandoned. A crossing settles output, error or abandoned. They are closed because a consumer that cannot rely on them cannot be written once and work everywhere. A host MUST NOT emit any other value in these attributes.

Four more sets are closed the same way and for the same reason: code_mode.observes_crossings, code_mode.crossing_edge, code_mode.crossing.timing, and the entries of code_mode.attested. A host MUST NOT emit a value outside them. A consumer that meets a value it does not know in one of the first three MUST treat that attribute as absent and apply the “absent reads as” rule in section 3; an unknown attested entry is ignored, which leaves the fields it would have upgraded at their baseline. Neither is a reason to discard the span.

These come from profiling nineteen code-mode implementations and adversarially testing every candidate rule against them. The surviving invariants are numbered C1 to C15 and X1 to X5 in Appendix A, and each new attribute below names the one it carries.

3. Declaration

Five attributes say what the host can and cannot see. Without them an absence of crossing spans has two readings a consumer cannot distinguish: the program made no calls, or the host cannot see the calls it made. That distinction is the difference between a trace you can reason from and a trace you cannot.

They are span attributes, on every span this document defines. Not Resource attributes, not instrumentation scope attributes. Section 3.2 says why, and the reason is not a preference.

AttributeTypeRequirementValues
code_mode.observes_crossingsstringRequiredall, some, none
code_mode.unmediated_egressbooleanRequiredtrue, false
code_mode.crossing_edgestringConditionally Required: when observes_crossings is not noneinvocation, dispatch
code_mode.attestedstring[]Requiredentries from the closed list in section 6.4; empty array when the host attests nothing
code_mode.attested_attributesstring[]Conditionally Required: with host_attributes, when the host measured any of its own attributeskeys in the host’s own namespace that the host itself measured, section 8
code_mode.relayed_attributesstring[]Conditionally Required: with host_attributes, when any of the host’s attributes came from a targetkeys the host passed through unchanged from a target, section 8

observes_crossings: all claims that every invocation routed through the host-provided surface is recorded. some claims the host mediates but records a subset, by policy or by mechanism. none says the host does not mediate calls at a call boundary. The value says nothing about whether other paths out of the program exist.

all is a claim about the whole path, not about the instrumented function. It is false if anything can answer the program without reaching the point the host records: a call-count cap, a deadline guard, a rate limiter, a cache, a permission check that refuses before dispatch. Such a refusal is an invocation the program made, and section 5.3 mints refused for exactly it, so a host that declares all either records those or is not all. The check is mechanical: exercise every refusal path and count the spans against the calls.

This is worth stating because it is the most damaging error available and the easiest to make. Every attested field is conditional on the declaration, and a host is usually instrumented at the function its bridge exposes, which on many implementations sits one layer beneath the guards that answer the program first. In a real integration of this specification a competent engineer declared all on such a host; four calls produced two spans, and the declaration said two was all of them.

unmediated_egress: true says the program has a way to reach the outside that the host does not see: raw network access, subprocess execution, an isolation layer that can be escaped. Consumers use it to refuse the inference “N crossing spans, therefore N external calls”.

crossing_edge says which edge a crossing span describes. invocation is what the program asked for at the call boundary, which is the program’s view. dispatch is what the host sent toward the target, recorded at the host’s own egress point, after any rewrite, retry decision or policy step. The two edges do not agree on cardinality: a retry or a refusal makes one invocation into zero or several dispatches, and a host that bundles invocations into one request makes several into one.

Absent reads as. A consumer that finds no declaration on a span reads observes_crossings as none, unmediated_egress as unknown and treats it as true, crossing_edge as unknown, and attested as empty. These are the weakest readings, and they are what a host gets for saying nothing.

3.1 Repetition

Every execution span and every crossing span carries the declaration. It is repeated per span, not stored once.

  • Every span a host emits for one dispatch MUST carry the same declaration values.
  • A host MUST compute those values from its own configuration for the profile it enforced, never from anything the program wrote or the caller claimed. A profile selected per dispatch is legitimate only when the host itself enforces the resulting capability, for example a network flag its own sandbox honours.
  • A host that cannot tell which profile served a dispatch MUST declare the weakest values that cover every profile it can reach, and MUST attest nothing.
  • A host MAY additionally place the same attributes on the Resource, for backends that facet on Resource. The span attributes are normative. Where the two disagree, the span wins and a consumer SHOULD count the disagreement.

A crossing span carries the declaration because a crossing span can reach a consumer without its execution span: sampled separately, exported in a different batch, or emitted for an execution the host never closed. A record that cannot be read alone is a record that has to arrive with its context intact, and nothing in a telemetry pipeline promises that.

The cost is five attributes on every span. The default AttributeCountLimit is 128 and a crossing span defined here carries at most fifteen attributes, so the repetition fits. It is not free: an execution with a thousand crossings pays for the declaration a thousand times, and no attribute interning is guaranteed on the wire.

3.2 Not Resource

An earlier draft put the declaration on the Resource. Two independent facts killed it, and either one alone is enough.

An API-only emitter cannot write a Resource attribute, ever. specification/resource/sdk.md: “A Resource is an immutable representation of the observed entity for which telemetry is being produced”, “Resources are immutable”, and “a resource can be associated with the TracerProvider when the TracerProvider is created. That association cannot be changed later.” specification/library-guidelines.md says instrumentation depends only on the API and cannot implement resource detection or instantiate providers. The Resource belongs to the application owner, and the application owner is the party that does not know about code mode. A host library can ship a resource detector and hope the owner wires it, which is a deployment step that will sometimes not happen, and silence reads as none.

Capabilities are not static per process. A profile MAY be selected per dispatch from a parameter the caller supplies, provided the host itself enforces the resulting capability. The profile survey has shipped examples. smolagents picks a local executor or a remote one per agent, in one Python library and one process, and those two executors do not have the same answer for observes_crossings. mcp-use picks a VM executor or a hosted sandbox per session, in one npm package, and those two do not have the same answer for attested or for crossing_edge. One TracerProvider has exactly one Resource, so a Resource cannot say that an absence of crossings on this execution means none happened while on that execution it means the host cannot see them. That inference is the only thing the declaration exists for.

Instrumentation scope attributes are not a fallback. The trace API has accepted scope attributes since specification 1.13.0, but the JavaScript API’s TracerOptions carries only schemaUrl, and the reference emitter for this model is JavaScript. A carrier that does not exist in the first language an implementer will reach for is not a carrier.

What the Resource loses, stated plainly: immutability. A Resource could not vary per span, and that was a structural guarantee the repetition rule now has to state as a rule instead. The two bullets in section 3.1 are what replaces it, and they are enforced by nothing but the host’s own code.

service.name and service.instance.id identify the host itself, on the Resource, as usual. This document mints no attribute for that.

4. Execution span

One dispatch of one program is one span.

Nameexecute_code {gen_ai.tool.name}, or execute_code when the dispatch did not arrive as a named tool call
KindSERVER when the dispatch arrived over a wire from an agent process; INTERNAL when the runtime dispatched it in-process
Startswhen the host first observes the dispatch, which for a host that runs the program is acceptance of the submission, including submissions it then rejects
Endswhen the host closes the execution
Parentthe caller’s span, taken from the incoming request context by the usual W3C propagation

gen_ai.operation.name is execute_code. That value is not in the well-known list upstream and this document proposes it. The nearest existing values, invoke_workflow and plan, both describe model-driven orchestration rather than the dispatch of one program, and execute_tool is taken by the crossing and by MCP’s own rule. The cost of a new value is real: no existing consumer recognises it until it lands upstream. The cost of reusing execute_tool is worse, because then an execution and the crossings inside it are indistinguishable by operation name, and any consumer counting tool calls counts each execution twice.

Exactly one span per dispatch carries code_mode.execution.disposition, and that span is the execution span. A host that already emits an MCP server span covering exactly this dispatch MAY put the execution attributes on that span instead of creating a child, following the MCP convention’s own rule against duplicating a span another instrumentation already owns. It then leaves gen_ai.operation.name and the span name as MCP sets them. A consumer finds the execution by the disposition attribute, not by the name. Section 17 records that two shapes is one too many.

4.1 Status

code_mode.execution.dispositionSpan status
completedleave Unset
failedError
terminatedError
abandonedleave Unset

Instrumentation does not set Ok. specification/trace/api.md: “Generally, Instrumentation Libraries SHOULD NOT set the status code to Ok, unless explicitly configured to do so.” The emitter this document describes is an instrumentation library, so completed is left Unset unless the application owner turns Ok on explicitly. The record format this replaces mapped completed to OK, which was legal there because a sink is an application. The emitter is not.

The collapse this produces, stated as a table because nobody should discover it from a dashboard:

Statuscovers
Unsetcompleted, abandoned, and every execution still running
Errorfailed, terminated

Four dispositions become two observable states, and three outcomes become two on the crossing span (section 5.1). In the one field every backend aggregates, alerts on and colours, a completed execution is indistinguishable from an abandoned one, and a failed execution is indistinguishable from a terminated one. C4 says that when an end exists the host can tell completed from every other disposition; after this move that claim lives only in the attribute.

So: code_mode.execution.disposition and code_mode.crossing.outcome carry the normative values. A consumer MUST read them. Span status is a display hint. That is why both attributes are Required.

terminated is Error even when the stop was an ordinary user cancellation. An operator who does not want cancellations in an error rate filters on code_mode.execution.disposition and error.type rather than on Status. The alternative considered was Unset for a terminated execution whose error class is cancelled, which was rejected because it makes a cancelled run indistinguishable from a clean one in Status, and because one rule with no exception is one fewer thing to get wrong.

The status description carries a closed-vocabulary value, not the host’s error message. Set it to error.type, or to the disposition. This diverges from the MCP convention, which sets the description from the error message, and the divergence is deliberate: the description is the one field on a span that has no attribute key, so nothing can carry a provenance label beside it. On most hosts an error message is the program’s own words, and putting those in the one unlabellable field publishes a program claim as prose that every UI renders as the reason a run failed. A closed-vocabulary host-observed value needs no label, which closes the problem by construction rather than documenting it. The message itself belongs in code_mode.error.body, where it can be labelled and where section 7 can say what was cut.

4.2 Attributes

AttributeTypeRequirementSource
gen_ai.operation.namestringRequiredthe constant execute_code
code_mode.execution.dispositionstringRequiredone of completed, failed, terminated, abandoned
the five declaration attributessection 3Requiredsection 3
code_mode.execution.idstringRequiredthe host’s own id for this execution; see below when it has none
code_mode.program.hashstringRecommendedsha256: and 64 lowercase hex digits over the UTF-8 bytes of the dispatched program text
code_mode.program.languagestringRecommendeda role hint: javascript, typescript, python, starlark. Omit rather than guess
gen_ai.tool.namestringConditionally Required: when the dispatch arrived as a named tool callthe name of the code-mode tool, for example execute
gen_ai.tool.call.idstringRecommendedthe caller’s tool-call id for this dispatch
gen_ai.conversation.idstringConditionally Required: when the host’s grouping is a conversation or agent sessionsee below
mcp.session.idstringConditionally Required: when the dispatch arrived over MCP in a sessionthe MCP session id
error.typestringConditionally Required: when the status is Errorsection 4.3
code_mode.program.textstringOpt-Inthe program text the host dispatched
gen_ai.tool.call.resultanyOpt-Inthe value the host returned on its return channel
code_mode.output.<channel>anyOpt-Inone per captured output channel: stdout, stderr, logs, files
code_mode.error.messagestringOpt-Inthe human-readable reason, which the status description no longer carries
code_mode.error.bodyanyOpt-Inthe raw error object as the host produced it
code_mode.captureanyRecommended: when the host shortened or removed any value above, or records sizessection 7

code_mode.execution.id earns its place because the span id is minted by the SDK and is not the id in the host’s own logs. An engineer holding a log line that names an execution needs a way back to the span.

It is Required rather than Recommended, and the difference matters more than it looks. A host that omits it produces spans that cannot be gathered into a run by any span-scoped query, and such a query returns no rows rather than an error, so the omission is invisible until someone is debugging at the wrong moment. A host that has no id of its own MUST mint one and use it on every span of that dispatch. A minted id answers “the crossings of this execution” exactly as well as a real one; what it cannot do is match a host log line, which is the reason a host that has an id should pass it rather than let one be made up.

code_mode.program.hash carries C1. It is how two dispatches of the same text are matched, and it is the only thing left when the program itself is withheld for privacy. A host that withholds the text emits the hash and records the withholding in code_mode.capture.

code_mode.program.text is Opt-In because it is agent-written code and routinely carries customer data. The program is not guaranteed to be everything that ran, nor byte-identical to what the runtime parsed, and a dispatch of several ordered segments may stop before the last one runs. A consumer MUST NOT infer from the disposition that any particular part of the program executed.

code_mode.program.language carries C2: a host may not know the language it runs, and the label exists for display and routing only. A consumer MUST NOT use it to predict what will parse. The core code.* registry was checked at 2026-09-19 and has no language attribute: it holds code.column.number, code.file.path, code.function.name, code.line.number and code.stacktrace as Stable, and five deprecated spellings. telemetry.sdk.language is the telemetry SDK’s own language, not the program’s.

gen_ai.conversation.id is the session grouping, and only when the grouping really is a conversation or agent session. A container id or a worker id is not a conversation: a host with one of those uses an attribute in its own namespace, per section 8. The upstream rule holds: when no identifier is available, do not populate it, and never fall back to a new UUID, a trace id or a hash of the request.

code_mode.output.<channel> carries X4. The channel name is the last component. Presence is the host’s declaration that it captures that channel, so a host that captures a channel emits it on every execution span, empty value and all. Without that rule an absent stderr means either “not captured” or “captured and empty”, and a consumer building an inventory from observed channels gets a different answer per run of the same host. Because the attribute is Opt-In, this rule binds only when the application owner has turned it on; when it is off, nothing is emitted and nothing is claimed.

4.3 Error type

error.type is Stable, low cardinality, with a well-known fallback value of _OTHER. Set it only when the status is Error. Recommended values, each a low-cardinality identifier:

validation for a rejection before the program ran, including a parse failure; runtime for an ordinary error the program raised after admission; timeout for the host’s own time limit; resource_limit for a declared non-time limit, such as memory or output size; cancelled for a stop by the caller or another external actor; approval_rejected for a gate that declined; host_failure for the host’s own process or infrastructure failing, and for a reconciliation that closed the record abandoned; _OTHER when the host has a disposition and no reason it can name.

A host MAY use a more specific low-cardinality identifier, such as an exception class name, and SHOULD document the values it emits. It MUST NOT put an unbounded value here: this attribute is one most backends facet on.

Read section 6 before trusting this attribute. On most hosts the class is derived from something the program wrote.

4.4 Missing states

A span is exported when it ends. A dispatch the host never closed is not in the trace at all, and is indistinguishable from one that never happened. The running state and the unresolved state have no representation here. That is the price of the move.

Closing a past-deadline execution abandoned at a later reconciliation is what puts it back in the trace, and it is the host’s job, never the consumer’s. A consumer MUST NOT synthesize an ending for a span it never received, and MUST NOT read the duration of an abandoned execution span as how long the program ran: the span ends when the host gave up.

A host SHOULD emit a log record when a dispatch starts, and MAY emit one when it ends. This is the only thing in the model that says a run is in flight right now, so it is what makes the running state visible at all.

event.namecode_mode.execution.started, or code_mode.execution.ended
SeverityINFO. A dispatch merely running is not something to page anyone about
Attributesthe execution span’s attributes, plus trace_id and span_id

The two ids are what make this a view of the trace rather than a second model. A consumer joins the records to the spans on them, and nothing here duplicates what a span already carries once the run has finished. A consumer MUST NOT depend on these records existing, because a host with no logging pipeline emits none.

5. Crossing span

One recorded invocation from the program across the host-provided boundary is one span.

Nameexecute_tool {gen_ai.tool.name}
KindCLIENT when the host forwards the call toward a remote target; INTERNAL when the host serves it in its own code
Startswhen the host observes the invocation
Endswhen the host determines the outcome, or when it closes the crossing unsettled
Parentthe execution span

The name and kind follow the existing gen_ai.execute_tool.internal span, which uses INTERNAL, and mcp.client, which uses CLIENT. Both precedents are upstream; this document borrows rather than invents. Span kind carries no information about whether the host mediated the boundary or merely parsed the program’s claim about it, so it is orthogonal to observes_crossings and to provenance.

Parentage places a crossing; it does not identify its execution. A crossing span is a child of its execution span, which is how a viewer nests them. Where the program runs in another process the host propagates the execution span’s context to the code that serves the bridge, so the crossing span is still a child. Section 6.6 says why that context must never come from inside the sandbox.

Parentage carries the parent’s span id, which is minted by the SDK and is not the id in the host’s own logs. So code_mode.execution.id is Recommended on a crossing span as well, for exactly the reason section 3.1 repeats the declaration: a crossing span reaches a consumer without its execution span often enough that it must be readable alone. Three cases, each of which was found by running the query rather than by reasoning about it:

  • A span-scoped query cannot join. A backend that filters spans one at a time, which is what TraceQL and every span-search box do, cannot express “the crossings of execution X” when only the parent carries X. A query written that way matches nothing, and it fails silently.
  • A log line carries the host’s execution id, not a span id. Going from an operator’s log line to the calls that run made needs the same key on both sides.
  • A process that dies mid-run orphans its crossings. The execution span was never exported, so the parent id resolves to nothing and the crossings are unreachable by any key an operator holds.

An earlier draft of this document argued that parentage made the attribute unnecessary. That was wrong, and it was wrong in a way that only showed up when someone wrote the query.

Crossings may overlap and carry no order. Sibling spans have no ordering in OpenTelemetry, which is exactly right: C10 says crossings within an execution are unordered unless seq or host-clock timestamps say otherwise. A consumer MUST NOT infer order from the order spans arrive, and MUST NOT assume crossings are sequential.

A crossing served by another execution is that execution’s span, as a child of the crossing span when the context propagated, and as a span link when it did not. This is C13, and OpenTelemetry satisfies it natively; nothing new is defined for it.

Open crossings when the execution ends. When an execution ends, the host ends every crossing span still open that it can still account for, with outcome abandoned, before it ends the execution span. A crossing whose bookkeeping was lost with the process that opened it produces no span. A settlement that arrives after the crossing was closed is not attributed to it.

A host that observes one SHOULD record it as a span event named code_mode.late_settlement, on the execution span, carrying gen_ai.tool.call.id to name the crossing it belongs to. Nothing is minted for that: it is the attribute the crossing span already carries for the host’s own id, and a host that did not set it there has no name to give the event either. Not on the crossing span: that span has ended by the time a late settlement exists, and the specification makes every operation on an ended span a no-op, so the event would be silently discarded. Where the execution span has ended too, which is the usual case because the execution ending is what closed the crossing, the trace has nowhere to put it and the log record in section 4.4 is the only place left. Section 16 records that.

5.1 Status

code_mode.crossing.outcomeSpan status
outputleave Unset
errorError
abandonedleave Unset

Three outcomes, two states. Unset covers output and abandoned, which are the two outcomes a reader most needs to tell apart. code_mode.crossing.outcome is Required for that reason, and section 4.1’s rule applies here too: the attribute is normative, Status is a display hint.

abandoned is not a failure and not a success. It says the host stopped observing and closed the record before it had determined either, usually because the execution ended first. It is not a claim that the target never responded.

Outcomes describe what the host itself determined at its own instrumentation point, not what the program observed. The outcome is fixed at the instant the host accepts the target’s answer or its own refusal. A fault in a later delivery step does not change it: if the host determined output and the value then failed to reach the program, the outcome stays output.

5.2 Attributes

AttributeTypeRequirementSource
gen_ai.operation.namestringRequiredthe constant execute_tool
gen_ai.tool.namestringRequiredthe target
code_mode.crossing.outcomestringRequiredone of output, error, abandoned
the five declaration attributessection 3Requiredsection 3
code_mode.crossing.timingstringConditionally Required: when the host synthesized either span timeone of start_only, end_only, none; section 5.4
code_mode.execution.idstringRequiredthe same id as the execution span this crossing belongs to
gen_ai.tool.call.idstringRecommendedthe host’s own id for this crossing
code_mode.crossing.dispatchedbooleanRecommendedwhether the host sent this call toward its target; false for a refusal it answered itself, or a cache hit
code_mode.crossing.seqintRecommended: when the host declares observes_crossings: all and has an initiation orderinitiation order within the execution, from 1
gen_ai.tool.typestringRecommendedfunction, extension or datastore, when the host knows
error.typestringConditionally Required: when the status is Errorsection 5.3
mcp.method.name, mcp.session.idstringConditionally Required: when the crossing went over MCP and this span is the only span for itsection 5.5
gen_ai.tool.call.argumentsanyOpt-Inthe input, fixed at initiation
gen_ai.tool.call.resultanyOpt-Inthe output, only under outcome output
code_mode.error.messagestringOpt-Inthe human-readable reason for this call’s failure
code_mode.error.bodyanyOpt-Inthe raw error object as the host or target produced it
code_mode.captureanyRecommended: as on the execution spansection 7

The target goes in gen_ai.tool.name, uncut. Nothing is minted for it, because that attribute already means the right thing and reusing it is what makes a code-mode crossing legible to a consumer that has never heard of code mode. The target is whatever the host uses to name what was invoked: a tool name, a namespace.method, a server and tool pair, a URL, a path. This document does not interpret it. Section 16 states the price of that reuse on a host that does not attest the target.

Span names must stay low cardinality. A host whose targets are unbounded, URLs for instance, uses a bounded form in the span name and keeps the full target in gen_ai.tool.name. This is the same allowance the MCP convention makes for resource URIs.

Input and output are typed any. Record them in structured form where the API supports it, and as a JSON string otherwise, which is the rule GenAI states for every any attribute. Both are Opt-In, because a crossing’s arguments and results are the most sensitive values in the trace. They are opaque: a consumer MAY display them and MUST NOT parse them, scan them for markers, or infer from them whether the value is inline text, base64 binary or a reference the host holds instead of content.

code_mode.crossing.seq carries C10 for hosts that have an order. A host that silently retries emits one span with the final outcome; the attempt count goes in the host’s own namespace, per section 8.

5.3 Error type

Recommended values: capability_error when the target returned an error for this call or the host cannot say more; validation when the input was rejected as malformed; refused when the host declined to dispatch by its own policy, before the target saw it; timeout for this call’s own time budget; cancelled when the target or host reported a cancellation; approval_rejected for a gate on this specific call; _OTHER when the host cannot classify.

Where the crossing was an MCP tool call that returned CallToolResult with isError true, error.type is tool_error, which is the MCP convention’s own rule.

An abandoned crossing carries no error.type. Nothing failed; the host stopped watching.

5.4 Times

A span always has a start and an end, therefore always a duration. OpenTelemetry has no representation for “settled, duration unknown.” This is a gap in the data model and nothing at this layer closes it.

Crossing times are optional in the model: their presence is the host’s declaration that it has host-clock timing for the crossing. A span needs both, so the host fills what is missing and says that it did.

host has starthost has end timespan startspan endcode_mode.crossing.timing
yesyesthe startthe endnot set
yesnothe startthe startstart_only
noyesthe endthe endend_only
nonothe host’s own reading at closethe samenone

The attribute is present exactly when a time was synthesized. All three synthesized cases are zero-duration spans, and a consumer MUST NOT read their duration as the crossing’s duration. Under none the span’s position says only when the host closed it.

An abandoned crossing usually lands in start_only. A trace viewer renders that as a zero-width tick inside its parent, with no status colour, which reads to the eye as a call that did not happen. It is the opposite. The case worth keeping in mind is a shipment the host started, stopped watching when the execution’s time limit fired, and closed abandoned. The goods may or may not have gone out. Section 15 shows what that trace looks like.

The inference this attribute exists to block, stated as a rule: a consumer MUST NOT treat an unresolved crossing as evidence the call is still in progress, or an abandoned crossing as evidence the target never responded. Both say only that the host stopped observing.

5.5 Over MCP

The MCP convention says: “If the MCP instrumentation can reliably detect that outer GenAI instrumentation is already tracing the tool execution, it SHOULD NOT create a separate span. Instead, it SHOULD add MCP-specific attributes to the existing tool execution span.”

A code-mode crossing dispatched over MCP is exactly that case. The host emits one span, carrying the crossing attributes and the MCP attributes together: mcp.method.name (tools/call for a tool call), mcp.session.id when there is a session, and mcp.resource.uri when the method takes one.

A sink reading a record after the fact cannot know whether a crossing was an MCP request, because no field says so. A host knows, and should say.

Where MCP client instrumentation is running in the same process and cannot detect the outer span, two spans per crossing is the result. That is a real failure and this document cannot prevent it from here; see section 16.

6. Provenance

6.1 The problem

A code-mode program is agent-written, and it can lie.

Some hosts build crossing records out of things the program wrote. Such a program can print a line that becomes a crossing span for a call it never made. It can throw an error whose message the host parses into an error class, choosing the classification the trace will show. It can return a value the host copies into an attribute. None of this requires an adversary: it is the ordinary consequence of building telemetry out of a channel the subject controls.

Consumers of these traces are increasingly language models. A model reading “the program deleted contact 42” needs to know whether the host saw that happen or the program said so.

OpenTelemetry has no notion of the host-observed versus program-determined distinction, and no per-attribute provenance channel anywhere in the data model. KeyValue is a key and a value. No field in Span, Resource, LogRecord, Link or Event annotates a value with where it came from. That narrow claim is the contribution. The broader claim, that OpenTelemetry has no trust or provenance work at all, is false and section 11 lists the open work that overlaps this.

6.2 Three classes

H, host-observed. Determined at a point the program cannot write through: the host’s own clock, its own id generation, an exit status, a call boundary the host mediates. Relative to the declaring host and conditional on its isolation not being bypassed. H means faithfully observed by this host. It does not mean true and it does not mean safe.

P, program-determined. Authored by the program, or computed by the host from a channel the program can write: the program text, standard output and error, thrown errors, files, return values, and anything derived from those.

T, target-relayed. Passed by the host unchanged from the target of a crossing, or produced by the host’s own handling of that crossing, such as a refusal or a policy error. The program did not shape it.

T does not mean the target saw the call. The name misleads on exactly the case an operator meets at three in the morning. A refusal the host answered itself is T, because the program did not shape it, and an operator reading a T error goes to the target’s own logs for a request that never left the process. code_mode.crossing.dispatched separates them. It is the host’s own knowledge, so it is H and carries no label, and a host that can tell SHOULD set it. This was found by someone working a real trace at a console, not by reading this table.

Whether a program is adversarial is a deployment question. These labels are about fidelity, not intent.

6.3 The design

A. A per-span attribute listing which attribute keys on this span are program-determined. Rejected. It fails open: an emitter that adds an attribute and forgets to add it to the list silently promotes a claim to an observation, and the failure is invisible. It also lets the list drift with whatever the emitter happened to write on that span, which is the wrong thing for the claim to track.

B. A naming convention that encodes provenance in the key, such as code_mode.claimed.tool.name beside code_mode.observed.tool.name. Rejected, and this is the decisive one. The attributes whose provenance is in question are the borrowed ones: gen_ai.tool.name, gen_ai.tool.call.arguments, gen_ai.tool.call.result, error.type. Encoding provenance in the key means minting parallel names inside gen_ai.* and error.*, which the naming rules forbid, or abandoning reuse and shipping a private vocabulary no existing consumer reads. That is the fragmentation this whole move exists to avoid. It has a second failure: the key changes when a host improves, so a host that moves from parsing stderr to mediating the boundary breaks every saved query, dashboard and alert built on its traces. A design that punishes the honest upgrade is wrong.

C. A declaration of what the host attests, from a closed list of model fields, plus a fixed table of baseline classes in this document. Chosen. It works on borrowed keys, because it names model fields and this document maps them to attribute keys once. It fails safe: a field nobody attested is a program claim, so an emitter that forgets something under-claims rather than over-claims. Its vocabulary is fixed by this document and cannot drift with the emitter’s attribute set, which is what separates it from A. And a consumer that ignores this scheme entirely still gets a valid, useful trace, with correct gen_ai.* attributes and correct parentage; it simply does not learn what was observed and what was claimed.

A, revisited, and adopted in part. The objection to A is about where the classes come from, not about the wire shape. A list the emitter assembles from whatever it happened to write can drift; a label the emitter computes from this document’s fixed table cannot, because the table is not the emitter’s to change and a field it has never heard of is simply not labelled, which under-claims. So C fixes the vocabulary and A’s shape carries it: alongside code_mode.attested, an emitter writes code_mode.provenance.<attribute key> next to every value whose effective class is P or T, and writes nothing beside a host-observed one. This is not a third design; it is C’s table, materialized. The record format’s own OpenTelemetry export has done exactly this since it was written, which is the standing proof that the shape is safe when the source is fixed.

Without it, a consumer learns nothing. code_mode.attested on its own is an answer to a question the consumer does not know to ask, and joining it against section 6.5 requires finding and reading this document. Almost nothing will. The label is what makes the idea survive contact with a backend, and it costs one attribute per unobserved field.

D. A span event or log record per value, carrying that value’s provenance. Rejected: unbounded volume, and it moves the claim away from the value.

The cost of C, stated plainly. A consumer must read one table in this document to know which attributes are provenance-bearing, and must read the declaration on the span. It cannot work that out from an attribute in isolation. And because the declaration is a span attribute rather than a Resource attribute (section 3.2), nothing structural stops a host varying it per span. Section 3.1 states the rule; only the host’s own code enforces it.

6.4 code_mode.attested

code_mode.attested is a string array from this closed list. A consumer MUST ignore an entry it does not know.

EntryUpgradesTo
crossing.targetgen_ai.tool.name, code_mode.crossing.seq and code_mode.crossing.outcome on crossing spansH
crossing.inputgen_ai.tool.call.arguments on crossing spansH
crossing.outputgen_ai.tool.call.result on crossing spansT
crossing.errorerror.type, the status description and code_mode.error.body on crossing spansT
execution.error.classerror.type and the status description on execution spansH
host_attributesthe attributes named in code_mode.attested_attributes (section 8)H
host_attributesand those named in code_mode.relayed_attributes (section 8)T

The entries are bundled rather than per attribute because they travel together. A host that observed the call boundary observed the target, the order and the outcome; a host that attested the outcome but not the target would be claiming something incoherent, and a closed list of six is shorter to write, shorter to read, and impossible to spell wrong.

Rules, and they are the rules the underlying model already states:

  • A host attests only what is true for every span it emits. There is no per-span attestation and no per-span opt-out. A host with both an observed path and a parsed path for the same field does not attest it.
  • A host that declares crossing_edge: dispatch SHOULD attest crossing.target. dispatch names what the host itself sent, so claiming that edge while not observing the target is contradictory. crossing_edge: invocation carries no such expectation: an invocation span is the program’s view by definition, which is what the next rule is for.
  • A host that derives crossings from program-written channels MAY still emit them. It simply does not attest them, and consumers read them as program claims. This is a feature. A stdout-parsing host that emits unattested crossings is more useful than one that emits nothing, as long as nobody mistakes the two.
  • Attestation makes a claim visible and attributable. It does not make it true. Nothing in a trace distinguishes a host reading its own call boundary from a host copying a value out of the program’s return and attesting it anyway. No format detects that. Attestation puts a name on the claim, which is all a format can do.

6.5 Baseline classes

Everything not in this table is H: span ids, parentage, start and end times, the five declaration attributes, code_mode.execution.id, code_mode.execution.disposition, code_mode.crossing.timing, code_mode.program.hash, gen_ai.tool.call.id, gen_ai.conversation.id, mcp.session.id, mcp.method.name, and every entry inside code_mode.capture.

AttributeSpanBaselineEntry that upgrades itAfter
code_mode.program.textexecutionPnoneP
code_mode.program.languageexecutionPnoneP
gen_ai.tool.call.resultexecutionPnoneP
code_mode.output.<channel>executionPnoneP
error.typeexecutionPexecution.error.classH
status descriptionexecutionPexecution.error.classH
code_mode.error.messageexecutionPnoneP
code_mode.error.bodyexecutionPnoneP
gen_ai.tool.namecrossingPcrossing.targetH
code_mode.crossing.seqcrossingPcrossing.targetH
code_mode.crossing.outcomecrossingPcrossing.targetH
gen_ai.tool.call.argumentscrossingPcrossing.inputH
gen_ai.tool.call.resultcrossingPcrossing.outputT
error.typecrossingPcrossing.errorT
status descriptioncrossingPcrossing.errorT
code_mode.error.messagecrossingPcrossing.errorT
code_mode.error.bodycrossingPcrossing.errorT
host’s own namespaceeitherPhost_attributes via attested_attributesH
host’s own namespaceeitherPhost_attributes via relayed_attributesT

These classes are written onto the span. For every attribute above whose effective class is P or T, an emitter writes code_mode.provenance.<that attribute's key> with the value P or T. A field that is host-observed, whether at baseline or after attestation, carries no such attribute, so absence means observed and a field the emitter has not heard of under-claims rather than over- claims. A host’s own attribute (section 8) is labelled P unless host_attributes is attested and the key is named in code_mode.attested_attributes.

Three readings worth spelling out, because they are the ones that surprise people.

A span’s own name and its status follow the attributes they came from. A crossing span is named from gen_ai.tool.name, so under an unattested host the span name itself is a program claim. The status follows code_mode.crossing.outcome, which follows crossing.target, so under an unattested host the red span in the UI is a program claim too. Neither the name nor the status can carry a label, which is limitation L3 in section 16.

The status description cannot be labelled. It has no attribute key, so nothing can name it in a list and nothing can carry its class beside it. It is in the table above so that a consumer knows what it is reading. A host SHOULD NOT put anything in the description that is not also in an attribute, and a host MUST NOT make the description the sole carrier of anything a consumer needs.

code_mode.program.text is never attestable, and neither is code_mode.program.hash’s subject. The program is by definition what the agent submitted. The hash is H because the host computed it, over content that is P. A host-observed hash of program-determined content is exactly what it sounds like, and it is still the right way to match two executions.

6.6 Rules

For any attribute whose effective class is P:

  1. Display it distinguishably. A viewer shows a program-reported marker or a distinct style. It does not show it the way it shows a timestamp.
  2. Exclude it from any aggregate presented as host-observed fact. Aggregates over program claims are legitimate when labelled as such. Section 9 is the hard form of this rule.
  3. When handing records to a language model, supply the classes alongside and state that P values are unverified program output.
  4. Never parse it. Do not scan a P value for structure or markers.

For T: the content came from the target, or from the host’s own handling of the crossing, and the program did not shape it. It is a claim about what the target returned, not something the host verified against the world.

For H: observed by the declaring host, subject to that host’s isolation. A consumer that does not trust the host trusts nothing.

And the rule for hosts, which is new here and has no counterpart in the record format, because trace context did not exist there:

A host MUST NOT accept trace context or spans minted inside the sandbox. If the program can supply a traceparent that the host then uses as the parent of its own spans, the program chooses where its execution appears in the trace, and can attach its records to another tenant’s trace. If the program can emit spans that reach the host’s exporter, it can write any attribute on any span, including the ones this document marks H. Context flows into the sandbox, never out of it. A host that gives the program its own tracer has made every span it emits program-determined, and must attest nothing.

7. Capture

OpenTelemetry has no way to say that a value on a record was shortened or removed. The SDK’s own attribute value length limit truncates silently. dropped_attributes_count says an attribute was dropped whole; it says nothing about a value that survived in shortened form. So this is minted, in one attribute.

code_mode.capture is a map from attribute key to a note about what the host did to that value. It is typed any: record it in structured form where the API supports it, and as a JSON string otherwise. Reading it is reading an envelope, not parsing a payload, so the rule in section 5.2 against parsing values does not apply to it.

FieldTypeMeaning
truncatedbooleanthe value on the span is a prefix of the host’s serialization of the original. It need not parse
redactedbooleanthe host removed or replaced the content by policy. The attribute itself may be absent
bytesintthe size in bytes of the host’s serialization of the original
hashstringsha256: and 64 lowercase hex digits over the host’s serialization of the original

Example, as a structured value: {"gen_ai.tool.call.result": {"truncated": true, "bytes": 6224, "hash": "sha256:..."}}.

Rules:

  • Every entry is H wherever it appears. Each is the emitter’s own record of what it did to its own capture, and no attestation entry moves it.
  • truncated and redacted are authoritative. Absence of an entry for a key means the emitter neither shortened nor removed that value.
  • Redacted is not the same as not recorded. An Opt-In attribute the host never records is simply absent, and its absence says nothing. redacted says the host held the value and removed it by policy. This is the only way to tell “we capture stderr and it is withheld here” from “we do not capture stderr”, now that presence alone cannot say it for an Opt-In attribute.
  • A host that drops a value because it could not serialize it, a cycle or a throwing serializer, dropped it by its own policy, and that is redacted. truncated is only ever a prefix of a serialization the host did produce.
  • bytes and hash describe the original, never the prefix, and are comparable only within one host, because serialization is host-defined.
  • A host that also writes an in-band marker into a value, such as a truncation note in the text, still sets truncated. The out-of-band flag is authoritative; the marker is content.
  • A program withheld for privacy is an absent code_mode.program.text, a present code_mode.program.hash, and an entry {"redacted": true, "bytes": N}.

Cut the value yourself. A host SHOULD shorten a large value in the emitter and record the truncation, rather than letting the SDK’s value length limit cut it silently downstream. Cut on a UTF-8 code point boundary. Set the cap below every downstream limit the host knows of, in the SDK, the collector and the backend.

7.1 Four rules

Each of these was a place two independent implementations of this document diverged, because it did not say. Each is now normative, and both agree.

A cap bounds the bytes that land on the span. For a payload written as JSON text that is the serialization; for code_mode.program.text, which is raw source, it is the raw UTF-8. Bounding the length of a JSON literal that attribute never becomes would spend a third of the allowance on escapes that are never written.

A value that cannot be serialized is redacted whole. JSON holds no NaN and no infinity. Writing null in their place turns a reading into a reading of nothing, which a reader who misses the flag takes at face value, so the emitter MUST drop the value rather than substitute. Such an entry carries neither bytes nor hash: both are defined over an original that could not be serialized.

A single value far past the cap is not read at all. An emitter MUST bound how much of one value it will read, and SHOULD set that bound at a small multiple of the cap. Serializing an enormous value is work a program can ask for without limit. A value refused this way is redacted, not truncated, because the emitter never produced a prefix of it.

An unpaired surrogate in the program text is replaced with U+FFFD, one per surrogate, before the hash is taken. Section 4.2 defines that hash over UTF-8 bytes and an unpaired surrogate has no UTF-8 encoding. This digest is the only key matching one dispatch to another across hosts, so it has to be the same number in every language.

And one thing that can never agree. bytes and hash are over the host’s own serialization, and two languages do not format numbers alike: an integral float, a negative zero and an integer past 2^53 all serialize differently. So these are a within-host key and never a cross-host one. Only code_mode.program.hash, which is over text rather than over a serialization, matches across hosts.

code_mode.capture is itself an attribute and is itself subject to those limits, so keep it small. It has one entry per value slot, so on the spans defined here it never exceeds a handful.

Section 16 records the two limits this does not close.

8. Host attributes

A host has values of its own: credits spent, a sandbox id, a subprocess exit code, an attempt count, a cache result, a model name. These go in the host’s own namespace, formed from its reverse domain name or its application name, for example com.acme.credits_used. They are never minted inside gen_ai.*, mcp.*, code_mode.* or otel.*.

They are P at baseline, like everything else the host did not declare it observed. Adding host_attributes to code_mode.attested is the gate, and it says only that the host is making provenance claims about its own attributes at all. Which claim, per key, comes from two lists:

  • code_mode.attested_attributes names keys the host measured itself, at a point the program cannot write through. Those are H and carry no label.
  • code_mode.relayed_attributes names keys the host passed through unchanged from a target, such as a credit count an API returned. Those are T.

A key in neither list stays P. A key in both MUST be refused: a host claiming both has not decided which claim it is making.

Two gates rather than one, for the reason the underlying model gives: a list alone is an upgrade path invisible in the attested declaration, which is the one place a consumer looks; and the entry alone would upgrade every host attribute at once, forcing a host to choose between recording a value the caller supplied and attesting its own meter.

Why T exists here. Without it the common case has no honest expression. A host that bills from a number its target reported cannot attest it, having not measured it, and leaving it P forbids the cost metric an operator needs, because section 9 bars a metric measured from a program claim. An earlier draft had only the one list, and the result was that every host with a real cost meter had an incentive to over-claim in a way nothing could detect. This is the same distinction sections 6.2 and 6.5 already draw for a crossing’s output; it was missing only here.

A host says what its own attributes mean, in code_mode.declared. A map from attribute key to { agg, unit, card, name }: whether the value can be summed (sum, last, none), what it counts, whether grouping by it is safe (low, high), and what to call it in a legend. It carries X5.

This is on the execution span only, and the difference from section 3.1 is the point. The capability declaration is repeated on every span because it changes how one span is read. This one is about combining values across spans, which is already a multi-span operation, so one carrier per trace is enough and a crossing does not pay for it.

It needs no attestation, because it is a claim about meaning rather than about fidelity. The two are separate and a consumer needs both: the declaration says a number adds up, and the provenance label says whose number it is. A value declared sum that carries a P label is still barred from a metric by section 9.

An earlier draft dropped this on the grounds that aggregation and unit belong to a metric instrument. That conflated two things. The instrument describes a metric the host chose to emit; this describes an attribute a consumer found, so that a consumer which has never heard of this host can do something correct with it. Nothing in OpenTelemetry carries per-attribute semantics, which is why it is minted here.

What reads it, honestly. No general-purpose backend does, and none will. A declaration is worth something to three consumers: a language model reading the trace, which increasingly is the consumer and which can act on it with no vendor support at all; a dashboard written against this specification, which can then render a host’s own fields without being rebuilt per host; and a collector deriving metrics, which can respect section 9 mechanically. Shipping it does not make Grafana understand your credit meter. It makes it possible for something to.

Use UCUM for units where one exists, By, ms, s, and a curly-brace annotation otherwise, {credit}, {token}.

9. Metrics

One rule. A metric point MUST NOT be keyed by, or measured from, any attribute whose effective class is P.

A metric point has no per-point provenance channel. A P value exported as a metric silently presents a program claim as fact, and a label added as a point attribute would become a cardinality dimension and make the metric unsummable across it. A P value stays on the span, where section 6.5 says what it is.

Four consequences, and the first is the one that costs something.

gen_ai.execute_tool.duration is conditional here. That histogram is Recommended upstream and its dimensions are gen_ai.tool.name, gen_ai.tool.type and error.type. On a code-mode crossing gen_ai.tool.name is the target, which is P unless the host attests crossing.target. So: a host MUST NOT emit gen_ai.execute_tool.duration for a code-mode crossing unless it attests crossing.target, and MUST NOT include error.type as a dimension unless it also attests crossing.error. A host that attests neither keeps the span and drops the metric.

MCP operation metrics are conditional for the same reason. mcp.client.operation.duration and mcp.server.operation.duration are keyed on mcp.method.name. On a crossing the host reconstructed from program output, the method name is the program’s claim, and the rule applies unchanged.

A host-specific value becomes a metric point only when host_attributes covers it (section 8).

What is always legal. The execution span’s own times, code_mode.execution.disposition and the five declaration attributes are H on every host. A duration histogram over executions, keyed on disposition, is therefore always sound.

9.1 Instruments

Two, both histograms, both in seconds.

InstrumentDimensions
code_mode.execution.durationcode_mode.execution.disposition, and error.type when present
code_mode.crossing.durationgen_ai.tool.name, code_mode.crossing.outcome and error.type, only when the host attests crossing.target

The condition on the second is the rule above, enforced rather than stated. On a host that did not attest the target, all three of those dimensions follow the target and are therefore the program’s words, so an emitter MUST drop them. What remains is an undimensioned duration distribution, which is worth little and is not a lie.

An emitter records nothing for a crossing that settled abandoned. Such a crossing is closed at its own start, so its duration is zero by construction, and recording it would put a fiction in the distribution.

This is the one place a host can enforce section 9 for itself. It cannot stop a span-metrics connector somebody else configured from deriving a metric out of a span name, which is why the collector configuration shipped alongside this specification exists.

Keep point attributes to low-cardinality dimensions. Never an execution id, a crossing id, a session id or anything per user.

Section 16 records the part of this rule that cannot be enforced from inside a host.

10. Integration

Two wrappers. That is the whole integration, and it is what both blind integrations converged on.

Wrapper one goes around the handler that runs the program. It starts the execution span before the program is dispatched, on the host’s own clock, extracting the incoming trace context from the request so the span has the caller’s span as parent. It ends the span when the host closes the execution, setting the disposition, and, on a non-normal end, the status and error.type. It ends every still-open crossing span first.

Wrapper two goes around the function the sandbox calls to reach the host. It starts a crossing span as a child of the execution span, records the target and the input at initiation, and ends the span when the host determines the outcome.

Three things determine whether the result is worth anything:

  • Wrapper two must be host code, outside the sandbox. A wrapper the program can reach, replace or observe is a channel the program writes through, and a host in that position attests nothing. This is the mechanical link between the integration shape and section 6: where the wrapper sits is what observes_crossings and code_mode.attested are describing.
  • Context flows in, never out. Wrapper one puts the execution span’s context where wrapper two can read it, in a context-local slot, in the bridge’s own per-execution state, or over the bridge’s own transport where the sandbox is another process. The program never supplies it. See section 6.6.
  • The emitter must not be able to change an outcome. No emitter fault may raise into the caller, and no metering call may relabel an outcome. Record the outcome first, then measure.

Both wrappers write the declaration attributes from section 3 onto every span they start. Neither wrapper needs an SDK. Both use the API only, so the application owner’s configured exporters receive the spans, which is the entire reason this is worth doing.

11. Upstream

State of open-telemetry/semantic-conventions-genai at 2026-09-19. Everything below is open, and everything below is Development.

UpstreamWhat it doesRelation to this document
PR #370gen_ai.attribution.link_type on span links: CAUSED_BY_GENERATION, RETRY_OF, INFORMED_BY. Titled “tool-call provenance”The same channel and an overlapping vocabulary for relating a retry or a replay to what it came from. The upstream version carries no count of attempts, which this project’s own earlier design argued is mandatory because there is no safe default
Issue #406Correlating GenAI spans with verified execution-environment attestation. States that “absence of attestation attributes must not be interpreted as a failed verification”The same fail-safe shape as code_mode.attested, for a different subject: it attests the environment, this attests what the host observed of the program
PR #445gen_ai.agent.paused, .checkpointed, .resumed, with gen_ai.agent.execution.id, pause.reason, resumed_from.type/.idSuspend and resume, which this project modelled as events and links. Its own text names its blocker: “LangGraph exposes no id spanning suspend and resume, so execution.id has no producer yet.” code_mode.execution.id is exactly that id, defined and Required by a model that has one
Issue #509Whether MCP tool calls are an execute_tool refinementDecides section 5.5
Issue #511Stabilizing inference and core agentic execution conventionsDecides when any of this can stop being Development
Issue #373Tool risk attributes for execute_tool and MCP tool call telemetryAdjacent. A risk label on a target the host did not observe has the same problem section 9 describes

What is not covered upstream, checked the same day: of the eleven gen_ai span types (inference.client, embeddings.client, retrieval.client, fetch_response.client, memory.client, create_agent.client, invoke_agent.client, invoke_agent.internal, execute_tool.internal, invoke_workflow.internal, plan.internal) none is a code execution or sandbox span. The MCP conventions model MCP at the JSON-RPC method level only. Nothing upstream models a submitted program, an execution, mediation, or the host-observed versus program-determined distinction.

12. Consumer rules

A consumer MAY rely on:

  • Every execution span carries exactly one code_mode.execution.disposition from the closed set, and every crossing span exactly one code_mode.crossing.outcome from its closed set.
  • Every span carries the declaration, and the declaration is the same on every span of one dispatch.
  • A crossing span’s parent is its execution span, or the two are joined by a span link.
  • Start and end times are the declaring host’s own clock readings, except where code_mode.crossing.timing says a time was synthesized. Crossing times and execution times are in the same clock domain, to whatever precision that host achieves.
  • code_mode.capture is authoritative about what the emitter did to a value.
  • Attributes covered by an entry in code_mode.attested are host-observed, or target-relayed, relative to the declaring host, and everything else in the section 6.5 table is a program claim.
  • Under observes_crossings: all, no invocation through the host-provided surface went unrecorded by that host, short of sampling and export loss.

A consumer MUST NOT:

  • Read Status as the outcome. Unset covers completed and abandoned. Error covers failed and terminated. The vocabulary attributes carry the answer.
  • Read a zero-duration crossing span as a crossing that took no time. Read code_mode.crossing.timing.
  • Treat the absence of crossing spans as evidence no calls happened. That inference needs observes_crossings: all, unmediated_egress: false, and a trace the sampler kept whole. Application owners SHOULD use a parent-based sampler so that an execution and its crossings are kept or dropped together; under a head sampler that decides per span, a missing crossing span means nothing at all.
  • Treat the absence of an execution span as evidence no execution happened. An execution the host never closed is never exported.
  • Infer order from the order spans arrive, or from ids, or assume crossings are sequential. Only code_mode.crossing.seq and host-clock timestamps carry order.
  • Assume the number of execution spans is the number of programs an agent submitted. A host emits one per dispatch, including dispatches it made on its own: a reactive re-run, a retry, a speculative branch, a shard of a data-parallel job.
  • Assume one crossing span is one dispatch to the target, or one invocation by the program. crossing_edge says which side the span describes, and nothing carries a count for the other side.
  • Read an abandoned crossing as evidence the target never responded, or an unfinished operation as evidence it is still running. Both say only that the host stopped observing.
  • Read the duration of an abandoned execution span as how long the program ran. The span ends when the host gave up, which is usually a later reconciliation.
  • Parse or interpret a payload value beyond displaying it.
  • Treat an attribute in the host’s own namespace as host-observed unless it is named in code_mode.attested_attributes and host_attributes is attested.
  • Use trace context, gen_ai.conversation.id, mcp.session.id, or gen_ai.tool.call.id on an execution span, for authorization or billing attribution. All of them are relayed from the caller, faithfully copied and unverified. gen_ai.tool.call.id on a crossing span is different: there the host minted it.
  • Assume the declaring host is trustworthy. Host-observed means observed by that host, relative to its own isolation.

13. Attributes

Twenty-one keys, one new enum value and one span event. Each names the invariant it carries. Everything else in this document reuses an attribute that already exists.

AttributeTypeWhereCarries
code_mode.observes_crossingsstringboth spansC5: mediation is declared, not assumed
code_mode.unmediated_egressbooleanboth spansC5
code_mode.crossing_edgestringboth spansX2: two edges
code_mode.attestedstring[]both spansX1: provenance
code_mode.attested_attributesstring[]both spansX1, for the host’s own attributes it measured
code_mode.relayed_attributesstring[]both spansX1, for the host’s own attributes a target reported
code_mode.declaredanyexecution spanX5: what the host’s own attributes mean, so a stranger can aggregate them
code_mode.execution.idstringboth spansC3: identity, and the only key that reads a crossing alone
code_mode.execution.dispositionstringexecution spanC3, C4: the closed disposition set Status cannot carry
code_mode.program.textstringexecution spanC1: one program per execution
code_mode.program.hashstringexecution spanC1: matching two dispatches of the same text
code_mode.program.languagestringexecution spanC2: language is a hint
code_mode.output.<channel>anyexecution spanX4: no universal output channel
code_mode.crossing.outcomestringcrossing spanC6: the closed outcome set Status cannot carry
code_mode.crossing.dispatchedbooleancrossing spanC6: whether an invocation reached a target at all
code_mode.crossing.seqintcrossing spanC10: no implicit order
code_mode.crossing.timingstringcrossing spanC7: crossing times exist only where the host observed them
code_mode.error.messagestringeitherC4: the reason, moved off the one field nothing can label
code_mode.error.bodyanyeitherC4: structured target errors that export would otherwise lose
code_mode.captureanyeitherC9: opaque payloads, truncation and redaction out of band
code_mode.provenance.<attribute>stringeitherX1: the class of the attribute it names, written only when that class is not host-observed

New value: gen_ai.operation.name = execute_code, for an execution span. Proposed, not upstream.

One span event: code_mode.late_settlement, on the execution span, for an outcome that arrived after the host closed the crossing (section 5). It names its crossing with gen_ai.tool.call.id and mints no key of its own. It carries C6: a crossing settles once.

Reused without change: gen_ai.operation.name, gen_ai.tool.name, gen_ai.tool.type, gen_ai.tool.call.id, gen_ai.tool.call.arguments, gen_ai.tool.call.result, gen_ai.conversation.id, mcp.session.id, mcp.method.name, mcp.resource.uri, error.type, and the whole trace data model: parentage, span links, span events and Status.

14. Dropped

For a reader who knew the retired JSON Lines format.

  • Id derivation. Span ids are minted by the SDK and trace context propagates natively. The host’s own execution id survives as code_mode.execution.id.
  • context.traceparent as a field. It is the incoming request context, extracted by the usual propagator. C13 is satisfied by parentage and links.
  • Start notices, and the unresolved state. A span exports when it ends. Section 4.4.
  • The supersede and conflict rules. They belong to a line-oriented format. Two spans with one span id are a backend problem, not this document’s.
  • dimensions. Restored as code_mode.declared in section 8, after being dropped on reasoning that conflated a metric instrument with an attribute a consumer found. Section 9’s rule, that a program-determined value never becomes a metric point, survives alongside it.
  • spec_version as a field. The scope version, and eventually schema_url, carry it.
  • Stateless sinks, malformed lines, line order. All properties of a JSON Lines stream.
  • OK for a completed execution. Section 4.1.

What did not drop: the model, the closed vocabularies, provenance, the capability declaration, and the two-wrapper integration shape.

15. Example

One execution, two crossings, the second abandoned. Every value below comes from a fixture this project has carried since before the move, and the emitter reproduces all of them.

The host is a synchronous bridge. It mediates every call at the call boundary and attests the target, the input and the output. The program cancels one order and then ships another. The shipment waits for an approval that never comes, the host’s own five minute limit fires, and the host closes the execution terminated. Before it does, it closes the shipping crossing abandoned with no end time, because it never determined an outcome for it.

14:00:00.000  host accepts the dispatch            execution span starts
14:00:00.080  program calls orders.cancel       crossing span starts
14:00:00.240  the delete returns                   crossing ends, outcome output
14:00:00.260  program calls orders.ship   crossing span starts
14:05:00.000  execution TTL elapsed                crossing closed abandoned, then execution ends

The execution span, as attributes:

AttributeValue
nameexecute_code execute
kindSERVER
statusError
gen_ai.operation.nameexecute_code
code_mode.execution.dispositionterminated
code_mode.observes_crossingsall
code_mode.unmediated_egressfalse
code_mode.crossing_edgeinvocation
code_mode.attested["crossing.target","crossing.input","crossing.output"]
code_mode.execution.id3c95e2578dd5e0169e81c566e43fac92
code_mode.program.hashsha256:bf15ddc985049f6ab5a1915a5e6235c149f48ef0e974318d2ca954a33ccec341
code_mode.program.languagejavascript
error.typetimeout

The abandoned crossing span, in OTLP/JSON. Note the equal start and end times, the timing attribute that says so, and the absent status, which is Unset. The span id is the SDK’s; the host’s own id for the crossing is in gen_ai.tool.call.id, which is the only place a reader holding a host log line can pick it up.

{
  "traceId": "0057b132ad41f0ea8a76f9299ba13793",
  "spanId": "7d1c04e9b8a3f265",
  "parentSpanId": "bb27b8faea63e97b",
  "name": "execute_tool orders.ship",
  "kind": 3,
  "startTimeUnixNano": "1789567200260000000",
  "endTimeUnixNano": "1789567200260000000",
  "attributes": [
    {"key": "gen_ai.operation.name",         "value": {"stringValue": "execute_tool"}},
    {"key": "gen_ai.tool.name",              "value": {"stringValue": "orders.ship"}},
    {"key": "code_mode.crossing.outcome",    "value": {"stringValue": "abandoned"}},
    {"key": "code_mode.crossing.timing",     "value": {"stringValue": "start_only"}},
    {"key": "code_mode.crossing.seq",        "value": {"intValue": "2"}},
    {"key": "code_mode.observes_crossings",  "value": {"stringValue": "all"}},
    {"key": "code_mode.unmediated_egress",   "value": {"boolValue": false}},
    {"key": "code_mode.crossing_edge",       "value": {"stringValue": "invocation"}},
    {"key": "code_mode.attested",            "value": {"arrayValue": {"values": [
      {"stringValue": "crossing.target"}, {"stringValue": "crossing.input"},
      {"stringValue": "crossing.output"}]}}},
    {"key": "gen_ai.tool.call.id",           "value": {"stringValue": "a1fa92d32b26e374"}},
    {"key": "code_mode.capture",             "value": {"stringValue":
      "{\"gen_ai.tool.call.arguments\":{\"bytes\":36,\"hash\":\"sha256:b8af7ef3554cb1900ad0506f62b274834ac850def21311f494d115e7a1d33843\"}}"}}
  ]
}

What a reader gets right from this span: the target and the input are attested, so the call was observed leaving, and the shipment really was requested. observes_crossings: all with unmediated_egress: false means there were exactly two crossings. The outcome is abandoned, so the host never learned whether it went out.

What a reader gets wrong if they read only the picture: a zero-width tick at 14:00:00.260, no status colour, inside a five minute parent. It looks like nothing happened. Section 16, L4.

16. Limitations

Each of these is a fact about what this design cannot do. None is closed by anything in this document, and none is argued away.

L1. The emitter cannot write a Resource, so the declaration is repeated on every span. Five attributes per span, on every crossing of every execution. There is no interning guarantee on the wire. The immutability a Resource would have given the declaration is replaced by a rule in section 3.1 that only the host’s own code enforces.

L2. Span status collapses four dispositions into two states, and three outcomes into two. completed and abandoned are both Unset. failed and terminated are both Error. Every default dashboard, alert and error rate in every backend reads that field and not the attributes. A consumer that wants the real answer must be told to read code_mode.execution.disposition and code_mode.crossing.outcome, and most consumers will not be.

L3. On a host that does not attest crossing.target, the crossing span’s NAME is a program claim, and standard pipelines read it as fact. The attribute beside it is now labelled, so a consumer that reads attributes can tell. The span’s own name cannot be labelled, and that is what span-metrics connectors, service maps and span-name-keyed alerting key on. Span-metrics connectors, service maps, span-name-keyed alerting and trace search all key on those two values and none of them reads the provenance table. Section 9 blocks the metric the upstream convention recommends, which is the part a host controls. It does not and cannot block a collector-side connector deriving metrics from span names. This is the price of reusing gen_ai.tool.name instead of minting a private key, and the reuse is still right, because a private key buys a consumer that reads nothing at all.

L4. There is no representation for “settled, duration unknown.” A span always has two times. An abandoned crossing, or any crossing on a host that does not record crossing times, becomes a zero-duration span that renders as a tick. code_mode.crossing.timing says so in an attribute, and no trace viewer reads it.

L5. A running or never-closed execution is not in the TRACE. It is indistinguishable there from a dispatch that never happened, because a span exports only when it ends. Section 4.4’s log record now covers it, so a host emitting all three signals can see work in flight. What remains is that the trace alone cannot, that a consumer reading only spans learns nothing, and that a host with no logging pipeline is back where it started. Closing a past-deadline execution abandoned at a later reconciliation is still the only way to put it in the trace itself.

L6. The SDK’s attribute value length limit truncates after the emitter has written its capture note. A value the host recorded whole can arrive shortened with nothing saying so. The default limit is infinite, so this bites only where an owner has set one, and the owner is the only party who can raise it. Not fixable from instrumentation.

L7. Sampling can drop part of a trace. Instrumentation cannot guarantee a whole trace is kept, so a missing crossing span can mean sampled rather than not observed. A parent-based sampler keeps an execution and its crossings together, and that is a recommendation to the application owner, not something the host can enforce. Not fixable from instrumentation.

L8. The attribute count limit, commonly 128, drops attributes past it silently. An execution span with many output channels and a large capture map can reach it. dropped_attributes_count says how many were lost, never which.

L9. any-typed attributes are not representable in the span attribute APIs of the three major languages. The specification’s AnyValue allows a nested map, but JavaScript’s SpanAttributeValue, Python’s types.AttributeValue and Java’s AttributeType are primitives and homogeneous arrays. In practice code_mode.capture, code_mode.error.body, code_mode.output.<channel>, gen_ai.tool.call.arguments and gen_ai.tool.call.result are JSON strings on a span today. Log records do not have this restriction.

L10. The status description has no attribute key, so nothing can label its provenance. Section 4.1 closes this by construction rather than by warning: the description carries error.type or the disposition, both closed vocabularies and both host-observed, so there is nothing there to label. The cost is a divergence from the MCP convention and the loss of a human-readable reason in the one place a UI shows it without being asked. The reason moves to code_mode.error.body, which is Opt-In, so a host that does not turn capture on has a less readable failure than it used to.

L11. Attestation is a claim, not proof. Nothing in a trace distinguishes a host reading its own call boundary from a host copying a value out of the program’s return and attesting it anyway. Detecting that needs a second observer in the path, under its own identity. No format detects it.

L12. gen_ai.operation.name = execute_code is not upstream. No existing consumer recognises it, and none will until a proposal lands. Until then a consumer filtering on known operation names does not see code-mode executions at all.

L13. Two spans per crossing are possible over MCP. The MCP convention’s anti-duplication rule is conditional on the MCP instrumentation detecting the outer span. Where it cannot, a crossing produces both a crossing span and an mcp.client span. Issue #509 is the place that gets decided.

L14. Everything this document reuses is Development. gen_ai.* and mcp.* carry no compatibility guarantee, and code_mode.* is a namespace this project owns and nobody else has agreed to.

L15. A closed vocabulary cannot be extended by a host, and that is both the point and the price. The host’s own namespace can add a field. It cannot add a value to code_mode.execution.disposition or code_mode.crossing.outcome, and it cannot make a consumer stop believing the closed one. The case that exposes this is a host whose logical run pauses at the end of one dispatch and resumes in a later one: each dispatch is its own execution, so the paused one reports completed, and a consumer following section 12 reads one logical run as having completed three times. A host attribute saying paused can sit right beside it and section 12 tells the consumer the closed value is normative. Closing the vocabularies is what lets a consumer be written once and work everywhere. It is also what makes the model unextensible at exactly these two points, and this document does not tell a reader which of its sets are open and which are closed in a way they could act on.

L16. Activating the execution span does nothing unless the application registered a context manager. An emitter can place its span in the active context, but with the API’s default NoopContextManager that call has no effect, and BasicTracerProvider.register() installs no manager. So neighbouring instrumentation inside a dispatch does not nest under the execution unless the application owner wired one, which NodeSDK does and a hand-assembled provider does not. The spans defined here are unaffected, because a crossing is given its parent explicitly.

L17. A late settlement usually has nowhere in the trace to go. A span that has ended takes no further events, by specification, and the thing that closed a crossing early is normally the execution ending, which closes the execution span too. So the one case the code_mode.late_settlement event was defined for is the case where no span is still open to carry it. What survives is the log record of section 4.4, which this document does not define. An abandoned crossing that did in fact settle is therefore, in the common case, indistinguishable in the trace from one that never did.

17. Open questions

Does this belong in open-telemetry/semantic-conventions-genai rather than here? That repository’s scope is the GenAI and MCP conventions; it was split out of open-telemetry/semantic-conventions, which now redirects there. Its stability level is Development throughout: none of gen_ai.* or mcp.* is Stable, and issue #511 is the open work to stabilize inference and core agentic execution. Its gaps are real, checked at 2026-09-19: no code execution or sandbox span among the eleven gen_ai span types, MCP modelled only at the JSON-RPC method level, and nothing anywhere that models a submitted program, mediation, or the host-observed versus program-determined distinction. Its naming rules leave exactly one route for an industry-wide attribute, which is a proposal to the specification.

The argument for upstream: every comparable project that kept its own vocabulary was eventually merged or demoted, and the failure was fragmentation rather than any technical flaw. The argument for here: this is unproven, code_mode.* is ours to change, and a rejected proposal is worse than no proposal. The way to settle it is not more argument. It is one pull request proposing the code-mode execution span, and what happens to it. Lead with code_mode.execution.id, because PR #445 names its absence as its own blocker and this model has the id it needs. Take the collision list in section 11 into that conversation rather than making the reviewer derive it.

Namespace. code_mode.* was chosen over the retired format’s mocon.*, because an upstream proposal named after a product would fail the naming rules, while gen_ai.* and mcp.* are named for their domain.

Merging with an existing MCP server span. Section 4 allows two shapes: a fresh execution span, or the execution attributes added to an MCP server span that already covers exactly the dispatch. Discovery is safe either way, because the execution span is defined as whichever span carries code_mode.execution.disposition. One rule would still be better than two, and picking one needs a test against a real MCP server instrumentation.

attested entry spelling. Entries name model fields (crossing.target) and section 6.4 maps them to attribute keys. Naming attribute keys directly would remove one lookup for a naive consumer, but it permits incoherent claims, attesting the outcome but not the target, and it is longer on the wire. The closed six-entry list was chosen. Reasonable people could pick the other one.

Standard code-mode metrics. Section 9 defines the prohibition and no instruments. An execution duration histogram keyed on disposition would be sound on every host (section 9), and a crossing count would not be on most. Whether that is a follow-up document is open.

The declaration on crossing spans. Section 3.1 requires all five on both span types, for a consumer that receives a crossing without its execution. If the repetition proves too expensive in practice, the smaller rule is code_mode.attested on crossings and the rest only on executions, because attested is the only one needed to read a crossing span’s own attributes. That would be two rules instead of one, which is why it is not the rule today.

The unresolved-execution log record. Section 4.4 says a host MAY emit one and defines nothing about it. If more than one host does it, two hosts will do it differently, and then it needs an event name, a body shape and a severity, which is a second document.

Appendix A. Invariants

These are the claims this specification is built on, not claims about any one implementation. They hold for a host as section 1 scopes one: a party that holds the program text it dispatched and can attribute the crossings it records to its own executions. Each attribute above cites the one it carries. They were derived by profiling nineteen implementations and adversarially testing every candidate against them, and they outlived the record format they were first written for.

  • C1. One program per execution. One execution is one dispatch of one program, never the session that contains it. The host holds that program text in full at dispatch. It is not guaranteed to be what an agent submitted for that dispatch, since a reactive runtime re-runs a dependent cell and a scheduler resumes a checkpoint, in which case the text is the host’s own. Nor is it everything that ran, nor what the runtime parsed.
  • C2. Language is a hint. A host may not know the language it runs. The label exists for display and routing only.
  • C3. Identity and disposition. Every execution has an id unique within its host and a host-observed start. If it ends, it ends with exactly one of completed, failed, terminated, abandoned. It may never end.
  • C4. Completed or not. When an end exists, the host can tell completed from every other disposition. Error detail is optional.
  • C5. Mediation is declared, not assumed. Whether the host observes crossings, and whether the program has a path out that the host does not see, differ by implementation and are declared.
  • C6. Crossing shape. Every recorded crossing has a target and an input fixed at initiation, and if it settles it settles as exactly one of output, error, abandoned. How many invocations or dispatches one record stands for follows from the declared edge (X2) and is not itself a core field.
  • C7. Host clock. Execution start and end are on the declaring host’s clock. Crossing times exist only where the host observed the crossing.
  • C8. Delivery varies. How an outcome reaches the caller, and whether the caller sees crossings, differ by implementation and are outside the contract.
  • C9. Opaque payloads. Inputs, outputs and results have no standard shape. Truncation and redaction are annotated out of band; consumers never parse values.
  • C10. No implicit order. Crossings within an execution are unordered unless seq or host-clock timestamps are present, and they may overlap.
  • C11. No universal session. Session, user and conversation identity are optional context the host passes through.
  • C12. Discovery is not universal. How the agent learns the callable surface is outside the contract.
  • C13. Nesting is a link. A crossing may be served by another execution. Correlation is by traceparent, not by a core field.
  • C14. Positions are not universal. Source positions for errors and crossings are not guaranteed and are not in core.
  • C15. Limits are not universal. Host-enforced limits and termination are not guaranteed.
  • X1. Provenance. Every field is host-observed, program-determined or target-relayed by a rule fixed in this specification. Only the attested list upgrades a field.
  • X2. Two edges. A crossing record describes either the program-facing invocation or the host’s dispatch toward the target. The host declares which.
  • X3. Environment is not fixed. The callable surface can change during an execution. Core does not record it.
  • X4. No universal output channel. Non-crossing outputs such as standard output are optional, per channel.
  • X5. Meaning is declared, identity is fixed. A host declares what its own attributes mean, so a consumer that has never heard of it can read them. No declaration reaches identity: not what an execution or a crossing is, not the closed dispositions and outcomes, not the reading of any attribute this document defines. This specification fixes the spine; everything above it is the host’s to declare.

Appendix B. Two implementations

This document has two independent implementations, in TypeScript and Python, and a harness that runs one scenario through both and diffs every attribute. That is the difference between a specification and a library with a document attached, and it is checkable rather than claimed.

They agree on everything this document defines: both spans, the capability declaration, every provenance label, the closed vocabularies, capture and its notes, crossing timing, seq, the MCP attributes, and both metric instruments with their section 9 gate.

Four things diverged when the harness was first run, in every case because this document was silent rather than because either implementation was wrong. All four are now stated in section 7.1 and both implementations follow them.

What remains is one language difference that no rule can close. JavaScript has a single number type, so 1.0 serializes as 1 where Python writes 1.0. Any bytes or hash over a payload therefore differs between the two, which is why section 7.1 marks them a within-host key. The program hash, taken over text rather than over a serialization, agrees.