Skip to content

trulens.core.dataset

trulens.core.dataset

Curation of recorded traces into persisted evaluation datasets.

Production traces already contain the examples a regression test needs, but add_ground_truth_to_dataset expects a dataframe that has already been extracted, joined, normalized and deduplicated. This module does that preparation, mapping a records dataframe onto GroundTruth rows through the existing dataset APIs.

Attributes

DEFAULT_BATCH_SIZE module-attribute

DEFAULT_BATCH_SIZE = 100

Number of ground truths written per database round trip.

ON_ERROR_RAISE module-attribute

ON_ERROR_RAISE = 'raise'

Fail on the first row that cannot be curated.

ON_ERROR_COLLECT module-attribute

ON_ERROR_COLLECT = 'collect'

Skip rows that cannot be curated and report them in the result.

PROVENANCE_COLUMNS module-attribute

PROVENANCE_COLUMNS = {
    "source_record_id": "record_id",
    "source_app_name": "app_name",
    "source_app_version": "app_version",
}

Metadata keys automatically preserved from a records dataframe.

Each is copied only when the source column is present and the caller has not already mapped that metadata key themselves.

Classes

TraceDatasetMapping

Bases: BaseModel

Mapping from records dataframe columns to ground truth fields.

Values are column names in the input dataframe, never expressions or callables, so a mapping is fully declarative and can be validated against the dataframe before anything is written.

Attributes
query class-attribute instance-attribute
query: str = 'input'

Column holding the query / input of the example.

query_id class-attribute instance-attribute
query_id: Optional[str] = 'record_id'

Column holding a stable identifier for the query, if any.

expected_response class-attribute instance-attribute
expected_response: Optional[str] = None

Column holding the corrected or expected response, if any.

expected_chunks class-attribute instance-attribute
expected_chunks: Optional[str] = None

Column holding the expected retrieval contexts, if any.

metadata class-attribute instance-attribute
metadata: Dict[str, str] = Field(default_factory=dict)

Metadata keys to preserve, mapped to the column they come from.

Functions
mapped_columns
mapped_columns() -> List[str]

Every dataframe column this mapping refers to, in a stable order.

missing_columns
missing_columns(dataframe: DataFrame) -> List[str]

The mapped columns that dataframe does not have.

CurationError

Bases: BaseModel

One input row that could not be turned into a ground truth.

Attributes
row_index class-attribute instance-attribute
row_index: Any = None

Index of the offending row in the input dataframe.

query_id class-attribute instance-attribute
query_id: Optional[str] = None

The row's query id, when one could be read.

reason instance-attribute
reason: str

Short, stable code for the kind of failure.

message instance-attribute
message: str

Human-readable detail.

CurationResult

Bases: BaseModel

Outcome of curating a records dataframe into a dataset.

Attributes
dataset_name instance-attribute
dataset_name: str

Name of the dataset written to.

dataset_id instance-attribute
dataset_id: DatasetID

Id of the dataset written to.

accepted class-attribute instance-attribute
accepted: int = 0

Rows that produced a distinct ground truth.

duplicates class-attribute instance-attribute
duplicates: int = 0

Rows that collapsed onto a ground truth already produced by this call.

Ground truth ids are content-addressed, so curating the same rows again in a later call also writes no new rows; that idempotency is not counted here because it is resolved by the database rather than by this call.

Note that a ground truth id covers its metadata too, so two records with the same question and answer but different provenance are two distinct ground truths. Pass include_provenance=False to deduplicate purely on example content.

rejected class-attribute instance-attribute
rejected: int = 0

Rows that could not be curated.

ground_truth_ids class-attribute instance-attribute
ground_truth_ids: List[GroundTruthID] = Field(
    default_factory=list
)

Ids of the ground truths written, in input order, without repeats.

errors class-attribute instance-attribute
errors: List[CurationError] = Field(default_factory=list)

One entry per rejected row. Always empty in raise mode.

processed property
processed: int

Total number of input rows considered.

Functions
errors_df
errors_df() -> DataFrame

Rejected rows as a dataframe, for inspection in a notebook.

CurationRowError

Bases: ValueError

A single row could not be curated.

Raised out of curate_records_to_dataset in raise mode; converted into a CurationError in collect mode.

Functions

curate_records_to_dataset

curate_records_to_dataset(
    dataset_name: str,
    records: DataFrame,
    db: Any,
    mapping: TraceDatasetMapping | None = None,
    expected_response_fn: (
        Callable[[Series], str | None] | None
    ) = None,
    dataset_metadata: dict[str, Any] | None = None,
    on_error: str = ON_ERROR_RAISE,
    batch_size: int = DEFAULT_BATCH_SIZE,
    include_provenance: bool = True,
    record_resolver: (
        Callable[[list[str]], DataFrame] | None
    ) = None,
) -> CurationResult

Curate a records dataframe into a persisted dataset.

See TruSession.curate_records_to_dataset for the user-facing entry point and argument documentation.

PARAMETER DESCRIPTION
dataset_name

Name of the dataset to write to.

TYPE: str

records

The rows to curate.

TYPE: DataFrame

db

Database to write through.

TYPE: Any

mapping

Column mapping. Defaults to TraceDatasetMapping().

TYPE: TraceDatasetMapping | None DEFAULT: None

expected_response_fn

Fallback for rows with no mapped correction.

TYPE: Callable[[Series], str | None] | None DEFAULT: None

dataset_metadata

Metadata for the dataset itself.

TYPE: dict[str, Any] | None DEFAULT: None

on_error

"raise" or "collect".

TYPE: str DEFAULT: ON_ERROR_RAISE

batch_size

Ground truths written per database round trip.

TYPE: int DEFAULT: DEFAULT_BATCH_SIZE

include_provenance

Whether to copy the source record id and app name/version into metadata.

TYPE: bool DEFAULT: True

record_resolver

Given record ids, returns their records. Used only when records is missing mapped columns.

TYPE: Callable[[list[str]], DataFrame] | None DEFAULT: None

RETURNS DESCRIPTION
CurationResult
RAISES DESCRIPTION
ValueError

If on_error is not a supported mode, batch_size is not positive, or a mapped column is missing from records.

CurationRowError

In raise mode, on the first unusable row.