Open the Traces dashboard
Where labeled turns land. Browse conversations and run Reflexes by hand; to export the raw turns, pull them from
GET /v1/reflex/traces.How the pieces fit
The rule that drives everything below: a turn is only classified if you wrap it in
begin(). Auto-instrumented LLM calls outside an interaction still get traced, but they have no event_id, so no label can attach.
Wire it up
1
Instrument your app
Install the SDK with the OpenTelemetry extra and initialize once at startup. After this, calls to instrumented SDKs are traced with no further changes.
2
Wrap the turns you want classified
Open an interaction with Keep
begin(), set the input, nest any tool calls, and close it with finish(). The event_id it mints is what every label links back to.convo_id stable across a conversation so the dashboard threads turns together, and user_id consistent so you can slice labels by user later.3
Turn on automatic classification
Pass Set a default for every turn by passing
evals on the turn. You choose which role each Reflex reads: user for the incoming message, assistant for the agent’s output. Morph classifies each role after the span lands, off your request path.evals to morph_tracing / morphTracing; a per-begin value overrides it. Omit both and nothing runs.4
Read the labels back
Open the Traces dashboard to browse turns with their labels, or pull them with the API. A freshly-traced turn shows “Classifying…” for a moment, then carries its
reflex_results.Python
GET /v1/reflex/traces returns LLM turns that carry text, newest first, each with the labels attached whether they ran from the SDK or by hand in the dashboard. Filter to one conversation with convo_id, page with limit / offset. A result’s selected names the winning class even when it’s the benign one (["Not Frustrated"], ["false"]), so test the label against the failure class, as above — non-empty selected is not a hit. See the field reference.Which Reflex on which role
Most safety and intent classifiers read the user’s message; response-quality ones read the agent’s output. A sensible starting set for a chat or coding agent:
Start with two or three that map to a failure you actually care about, watch the dashboard for a day, then add more. Custom Reflexes you’ve trained drop into the same
evals arrays by name.
Worked example: a GLM-5.2 agent, end to end
One Morph key runs the whole thing. The agent itself runs on GLM-5.2 (morph-glm52-744b) through Morph’s OpenAI-compatible endpoint; the same SDK call that points the OpenAI client at Morph also gets it auto-instrumented by tracing, so every model call is a span. Wrap each turn in begin() with evals and Reflexes label it off the request path. No second provider, no judge in the loop.
What to do with the labels
A label is only worth collecting if you act on it.- Alert. Poll
GET /v1/reflex/traces(or wire the dashboard) and page on-call whenjailbreakorguardrailfires, or whenuser-frustratedcrosses a rate you set for a conversation. - Build training sets. Filter traces by label to pull the exact turns you want — every
stuck-in-a-loopturn, every frustrated exchange — and feed them into evals or fine-tuning. The list endpoint returnsinput_textandoutput_textdirectly. - Track trends. Watch a label’s rate over time to know whether a prompt change actually reduced frustration or just moved it.
Classification rides the trace export, so it adds nothing to your latency. Under the hood, evals go through the same queue as the async batch API and are billed at the batch rate — 0.00025 past 1M/month. You pay only for the turns you put in
evals, not every traced span. See Reflex pricing.Backfill traces you already have
Turning onevals only labels turns going forward. To classify a backlog — every conversation from last month, scanned for jailbreaks and loops — run it from the Traces dashboard: select conversations, pick the Reflexes, and the labels land back on each trace. Under the hood that’s an asynchronous batch over the text of each turn, billed at the discounted batch rate. No code required.
Next steps
Tracing reference
Every config field, direct OTLP ingest, and the
/v1/reflex/traces schema.Reflexes overview
The nine default classifiers, response shape, and realtime
/predict.Train a Custom Reflex
When the defaults don’t match your failure modes, train one in ~30s and drop it into
evals.Batch classification
Label a backlog of up to 10,000 rows offline at the discounted rate.