trulens.core.dataset¶
trulens.core.dataset
¶
Curation of recorded traces into persisted evaluation datasets.
Production traces already contain the examples a regression test needs, but add_ground_truth_to_dataset expects a dataframe that has already been extracted, joined, normalized and deduplicated. This module does that preparation, mapping a records dataframe onto GroundTruth rows through the existing dataset APIs.
Attributes¶
DEFAULT_BATCH_SIZE
module-attribute
¶
DEFAULT_BATCH_SIZE = 100
Number of ground truths written per database round trip.
ON_ERROR_RAISE
module-attribute
¶
ON_ERROR_RAISE = 'raise'
Fail on the first row that cannot be curated.
ON_ERROR_COLLECT
module-attribute
¶
ON_ERROR_COLLECT = 'collect'
Skip rows that cannot be curated and report them in the result.
PROVENANCE_COLUMNS
module-attribute
¶
PROVENANCE_COLUMNS = {
"source_record_id": "record_id",
"source_app_name": "app_name",
"source_app_version": "app_version",
}
Metadata keys automatically preserved from a records dataframe.
Each is copied only when the source column is present and the caller has not already mapped that metadata key themselves.
Classes¶
TraceDatasetMapping
¶
Bases: BaseModel
Mapping from records dataframe columns to ground truth fields.
Values are column names in the input dataframe, never expressions or callables, so a mapping is fully declarative and can be validated against the dataframe before anything is written.
Attributes¶
query
class-attribute
instance-attribute
¶
query: str = 'input'
Column holding the query / input of the example.
query_id
class-attribute
instance-attribute
¶
Column holding a stable identifier for the query, if any.
expected_response
class-attribute
instance-attribute
¶
Column holding the corrected or expected response, if any.
expected_chunks
class-attribute
instance-attribute
¶
Column holding the expected retrieval contexts, if any.
metadata
class-attribute
instance-attribute
¶
Metadata keys to preserve, mapped to the column they come from.
Functions¶
mapped_columns
¶
Every dataframe column this mapping refers to, in a stable order.
CurationError
¶
CurationResult
¶
Bases: BaseModel
Outcome of curating a records dataframe into a dataset.
Attributes¶
accepted
class-attribute
instance-attribute
¶
accepted: int = 0
Rows that produced a distinct ground truth.
duplicates
class-attribute
instance-attribute
¶
duplicates: int = 0
Rows that collapsed onto a ground truth already produced by this call.
Ground truth ids are content-addressed, so curating the same rows again in a later call also writes no new rows; that idempotency is not counted here because it is resolved by the database rather than by this call.
Note that a ground truth id covers its metadata too, so two records with
the same question and answer but different provenance are two distinct
ground truths. Pass include_provenance=False to deduplicate purely on
example content.
ground_truth_ids
class-attribute
instance-attribute
¶
ground_truth_ids: List[GroundTruthID] = Field(
default_factory=list
)
Ids of the ground truths written, in input order, without repeats.
errors
class-attribute
instance-attribute
¶
errors: List[CurationError] = Field(default_factory=list)
One entry per rejected row. Always empty in raise mode.
Functions¶
CurationRowError
¶
Bases: ValueError
A single row could not be curated.
Raised out of curate_records_to_dataset in raise mode; converted into a
CurationError in collect mode.
Functions¶
curate_records_to_dataset
¶
curate_records_to_dataset(
dataset_name: str,
records: DataFrame,
db: Any,
mapping: TraceDatasetMapping | None = None,
expected_response_fn: (
Callable[[Series], str | None] | None
) = None,
dataset_metadata: dict[str, Any] | None = None,
on_error: str = ON_ERROR_RAISE,
batch_size: int = DEFAULT_BATCH_SIZE,
include_provenance: bool = True,
record_resolver: (
Callable[[list[str]], DataFrame] | None
) = None,
) -> CurationResult
Curate a records dataframe into a persisted dataset.
See TruSession.curate_records_to_dataset for the user-facing entry point and argument documentation.
| PARAMETER | DESCRIPTION |
|---|---|
dataset_name
|
Name of the dataset to write to.
TYPE:
|
records
|
The rows to curate.
TYPE:
|
db
|
Database to write through.
TYPE:
|
mapping
|
Column mapping. Defaults to
TYPE:
|
expected_response_fn
|
Fallback for rows with no mapped correction. |
dataset_metadata
|
Metadata for the dataset itself. |
on_error
|
TYPE:
|
batch_size
|
Ground truths written per database round trip.
TYPE:
|
include_provenance
|
Whether to copy the source record id and app name/version into metadata.
TYPE:
|
record_resolver
|
Given record ids, returns their records. Used only
when |
| RETURNS | DESCRIPTION |
|---|---|
CurationResult
|
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
If |
CurationRowError
|
In |