Selecting OTel GenAI Semantic Convention Attributes in TruLens¶
TruLens emits OpenTelemetry GenAI semantic convention
attributes alongside its own ai.observability.* attributes, so a metric can be
pointed at a gen_ai.* key directly, without TruLens-specific attribute
knowledge.
Every metric below is a built-in TruLens evaluator — the only thing that changes
is the selector pointing it at a gen_ai.* attribute:
| Scope | Metric here |
|---|---|
| A retrieval span's documents | Context Relevance, Groundedness |
| The generation span's message events | Answer Relevance |
| The whole trace | Logical Consistency |
| The whole conversation | Conversation Helpfulness |
For a worked coding-agent example, evaluating a Claude Code session assembled
from client hooks, see coding_agent_trace_evaluation.ipynb.
What TruLens emits, and what it does not¶
@instrument() sets gen_ai.* attributes automatically based on span_type.
You do not call any GenAI-specific API — you set the TruLens attributes as
usual and the gen_ai.* mirror is written for you.
span_type |
gen_ai.* span attributes emitted |
|---|---|
GENERATION |
gen_ai.operation.name, gen_ai.request.model, gen_ai.request.temperature, gen_ai.system, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens |
RETRIEVAL |
gen_ai.retrieval.query.text, gen_ai.retrieval.documents |
TOOL, MCP |
gen_ai.tool.name, gen_ai.tool.call.arguments, gen_ai.tool.call.result |
Message content is handled separately from metadata. Prompts and completions are
emitted as a gen_ai.client.inference.operation.details span event rather than
as span attributes, gated behind TRULENS_OTEL_CAPTURE_CONTENT for PII safety.
Event attributes live in their own container, so they have their own selector
field — span_event_attribute, optionally narrowed by span_event_name:
| Event attribute | Holds |
|---|---|
gen_ai.input.messages |
Structured input messages, with roles and typed parts |
gen_ai.output.messages |
Structured output messages |
Answer Relevance below is fed from those two.
One rule governs what becomes selectable at all: a span is persisted only if it
carries ai.observability.app_name. TruLens' span processor copies that from
context baggage onto every span started inside a recording context — including
spans it did not create.
So spans from third-party instrumentation such as openinference or
opentelemetry-instrumentation-openai are kept when they are emitted during a
with recorder: block. They arrive with span_type unknown, keep their own
attributes, and can be selected by span_name or with a
span_attributes_processor. Spans started outside any recording context have no
app_name in scope and are dropped at export.
gen_ai.retrieval.query.text and gen_ai.retrieval.documents are TruLens
aliases under the gen_ai namespace, not official OTel GenAI attributes — the
spec has no retrieval attributes yet.
Setup¶
# !pip install trulens trulens-providers-cortex snowflake-snowpark-python
import os
# Message content rides a span event, which is only emitted when content capture
# is opted in. Without this, the span-event selector below finds nothing.
os.environ["TRULENS_OTEL_CAPTURE_CONTENT"] = "true"
from snowflake.cortex import complete
from snowflake.snowpark import Session
from trulens.apps.app import TruApp
from trulens.core import Metric
from trulens.core import TruSession
from trulens.core.database.connector.default import DefaultDBConnector
from trulens.core.feedback.selector import Selector
from trulens.core.otel.instrument import instrument
from trulens.otel.semconv.trace import GenAIAttributes
from trulens.otel.semconv.trace import GenAIEvents
from trulens.otel.semconv.trace import SpanAttributes
from trulens.providers.cortex import Cortex
Connect to Snowflake. APP_MODEL generates answers; JUDGE_MODEL runs the
built-in evaluators.
snowpark_session = Session.builder.config(
"connection_name", os.environ.get("SNOWFLAKE_CONNECTION_NAME", "default")
).create()
APP_MODEL = "llama3.3-70b"
JUDGE_MODEL = "claude-sonnet-4-5"
provider = Cortex(snowpark_session=snowpark_session, model_engine=JUDGE_MODEL)
session = TruSession(
connector=DefaultDBConnector(database_url="sqlite:///genai_semconv_demo.sqlite")
)
session.reset_database()
An instrumented support agent¶
Three instrumented methods, one span type each. Note that nothing here mentions
gen_ai — these are the ordinary TruLens attributes, and the gen_ai.* mirror
is derived from them. The GENERATION span picks up gen_ai.request.model from
COST.MODEL, plus temperature, provider_name, and operation_name, which
the instrumentor looks for by name.
KB = [
"Password resets are self-serve from Account -> Security -> Reset password.",
"A reset link stays valid for 30 minutes, then must be requested again.",
"Resetting a password signs the account out of every active session.",
"Refunds for annual plans are prorated from the cancellation date.",
"Our headquarters is in Bellevue, Washington.",
]
class SupportAgent:
@instrument(
span_type=SpanAttributes.SpanType.RETRIEVAL,
attributes={
SpanAttributes.RETRIEVAL.QUERY_TEXT: "query",
SpanAttributes.RETRIEVAL.RETRIEVED_CONTEXTS: "return",
},
)
def retrieve(self, query: str) -> list:
"""Keyword overlap stand-in for a vector search."""
terms = {t for t in query.lower().split() if len(t) > 3}
scored = [(len(terms & set(doc.lower().split())), doc) for doc in KB]
scored.sort(key=lambda pair: pair[0], reverse=True)
return [doc for _, doc in scored[:3]]
@instrument(
span_type=SpanAttributes.SpanType.GENERATION,
attributes=lambda ret, exception, *args, **kwargs: {
SpanAttributes.COST.MODEL: APP_MODEL,
# Plain keys the instrumentor maps into gen_ai.* by name.
"temperature": 0.0,
"provider_name": "snowflake.cortex",
"operation_name": "chat",
# These become the gen_ai.input.messages / gen_ai.output.messages
# attributes on the span's inference event.
"prompt": kwargs.get("query"),
"completion": ret,
},
)
def generate(self, query: str, contexts: list) -> str:
prompt = (
"Answer the customer's question using only the context below. "
"Be concise.\n\nContext:\n"
+ "\n".join(f"- {c}" for c in contexts)
+ f"\n\nQuestion: {query}"
)
return complete(
APP_MODEL, [{"role": "user", "content": prompt}], session=snowpark_session
).strip()
@instrument(
span_type=SpanAttributes.SpanType.RECORD_ROOT,
attributes={
SpanAttributes.RECORD_ROOT.INPUT: "query",
SpanAttributes.RECORD_ROOT.OUTPUT: "return",
},
)
def answer(self, query: str) -> str:
return self.generate(query=query, contexts=self.retrieve(query))
Metrics¶
Four built-in evaluators, each pointed at a gen_ai.* attribute by its selector.
span_attribute takes the attribute name directly, so any gen_ai.* key TruLens
persisted can be selected — a scalar value, or several combined:
# Any single gen_ai.* key.
Selector(
span_type=SpanAttributes.SpanType.GENERATION,
span_attribute=GenAIAttributes.REQUEST.MODEL,
)
# Or reshape several of them before the metric sees them.
Selector(
span_type=SpanAttributes.SpanType.GENERATION,
span_attributes_processor=lambda attrs: attrs.get(
GenAIAttributes.USAGE.INPUT_TOKENS, 0
),
)
Context Relevance and Groundedness both read gen_ai.retrieval.documents, the
first scoring each document and the second checking the answer against all of
them together. trace_level=True hands the metric a Trace covering every span
in the record instead of one span's attribute, and must be the only selector on
that metric. .on_conversation() groups every RECORD_ROOT sharing a
conversation_id, ordered by start time, and attaches its score to the
conversation's last record.
m_context_relevance = Metric(
implementation=provider.context_relevance_with_cot_reasons,
name="Context Relevance",
).on({
"question": Selector(
span_type=SpanAttributes.SpanType.RETRIEVAL,
span_attribute=GenAIAttributes.RETRIEVAL.QUERY_TEXT,
),
"context": Selector(
span_type=SpanAttributes.SpanType.RETRIEVAL,
span_attribute=GenAIAttributes.RETRIEVAL.DOCUMENTS,
collect_list=False, # one LLM call per retrieved document
),
})
m_groundedness = Metric(
implementation=provider.groundedness_measure_with_cot_reasons,
name="Groundedness",
).on({
"source": Selector(
span_type=SpanAttributes.SpanType.RETRIEVAL,
span_attribute=GenAIAttributes.RETRIEVAL.DOCUMENTS,
collect_list=True, # one LLM call against all documents
),
"statement": Selector.select_record_output(),
})
m_answer_relevance = Metric(
implementation=provider.relevance_with_cot_reasons,
name="Answer Relevance",
).on({
"prompt": Selector(
span_type=SpanAttributes.SpanType.GENERATION,
span_event_attribute=GenAIEvents.EventAttributes.INPUT_MESSAGES,
span_event_name=GenAIEvents.CLIENT_INFERENCE_OPERATION_DETAILS,
),
"response": Selector(
span_type=SpanAttributes.SpanType.GENERATION,
span_event_attribute=GenAIEvents.EventAttributes.OUTPUT_MESSAGES,
span_event_name=GenAIEvents.CLIENT_INFERENCE_OPERATION_DETAILS,
),
})
m_logical_consistency = Metric(
implementation=provider.logical_consistency_with_cot_reasons,
name="Logical Consistency",
).on({"trace": Selector(trace_level=True)})
m_conversation_helpfulness = Metric(
implementation=provider.conversation_helpfulness_with_cot_reasons,
name="Conversation Helpfulness",
).on_conversation()
Record a conversation¶
Passing conversation_id to the recording context is what makes the
conversation-level metric possible.
agent = SupportAgent()
recorder = TruApp(
agent,
app_name="GenAI Semconv Selection",
app_version=APP_MODEL,
main_method=agent.answer,
feedbacks=[
m_context_relevance,
m_groundedness,
m_answer_relevance,
m_logical_consistency,
m_conversation_helpfulness,
],
)
turns = [
"How do I reset my password?",
"How long does that reset link stay valid?",
"Will resetting sign me out everywhere?",
]
with recorder(conversation_id="support-thread-1") as recording:
for turn in turns:
print(f"Q: {turn}")
print(f"A: {agent.answer(turn)}\n")
session.force_flush()
Compute and inspect the metrics¶
recorder.compute_feedbacks(raise_error_on_no_feedbacks_computed=False)
session.force_flush()
records, metric_names = session.get_records_and_feedback(
app_name="GenAI Semconv Selection"
)
present = [name for name in metric_names if name in records.columns]
records[["input"] + present].round(3)
The _with_cot_reasons evaluators persist the judge's reasoning alongside the
score, under <metric>_calls[...]["meta"]["explanation"]. That is where to look
when a score is surprising — the reasoning below is the metric's own account of
why it scored what it did.
# Each metric is shown on its own weakest record. Conversation-level metrics only
# attach to a conversation's last record, so they score on different rows than
# record-level ones.
for name in present:
scored = records[
records[f"{name}_calls"].apply(
lambda calls: isinstance(calls, list) and len(calls) > 0
)
].dropna(subset=[name])
if scored.empty:
continue
row = scored[name].idxmin()
calls = records[f"{name}_calls"].loc[row]
print("=" * 70)
print(name, "=", records[name].loc[row])
print("record:", records["input"].loc[row])
print("=" * 70)
print(calls[0].get("meta", {}).get("explanation"))
print()
Dashboard¶
from trulens.dashboard import run_dashboard
run_dashboard(session)