Evaluate AI applications and agents with Deepchecks’ custom judges, version comparisons, trace analysis and production quality monitoring.
Deepchecks
Explore features, practical uses and pricing below.
Deepchecks evaluates and monitors AI applications, including question-answering assistants, retrieval workflows and multi-step agents. Teams send application interactions or traces to the platform, define the qualities that matter, and inspect scores, labels and operational measurements. Version comparison helps examine a prompt or model change against an earlier implementation, while production monitoring tracks how the deployed system behaves over time.
The current Deepchecks website focuses on LLM Evaluation and Know Your Agent, its agent testing workflow. This is useful for engineering, AI quality and product teams that already have an application to assess. The platform organizes evidence for improvement; the team still decides which errors are acceptable, what a release must satisfy and how to respond when behavior changes.
A Deepchecks application represents a distinct AI system or workflow. Keeping separate applications for a support assistant and an internal research agent makes their evaluation rules easier to understand. A support answer might need to follow an escalation policy, while a research response might need evidence for each conclusion. Combining them under one broad definition of quality would hide meaningful differences.
The Python SDK quickstart shows creating an application, logging interactions under a named version and identifying the evaluation environment. Interactions include input and output, and additional fields can provide timing or token information. Deepchecks then calculates properties and estimated annotations asynchronously.
Instrument the information needed to answer a diagnostic question. An output alone may reveal an incorrect answer but not whether retrieval failed or the model ignored a correct source. Missing timestamps prevent useful timing analysis. Inconsistent version names can make an experiment difficult to reconstruct, even when every individual response has been uploaded successfully.
Start with a small group of representative records and check that the displayed fields match the application. Include a clear success, a known failure and an ambiguous case. This checks the integration and establishes whether the initial quality signals resemble the team's expectations before a larger upload.
Properties are measurements attached to interactions. The SDK guide gives examples such as grounding in context, avoided answers, fluency and sentiment. Different measurements describe different aspects of an output. A fluent answer can still be unsupported, and a refusal can be appropriate even when an avoided-answer signal is triggered.
Deepchecks also provides prompt properties for custom evaluation using an LLM as a judge. Users write guidelines and choose the data fields the evaluator sees. Numerical properties produce a graded score with reasoning; categorical properties classify an interaction into defined labels. The documentation describes testing a property on a sample before applying it broadly.
Write a rubric that a reviewer could apply consistently. For a shipping-policy assistant, specify what counts as a complete answer, which exceptions matter and when the response should ask for missing information. A vague instruction to evaluate accuracy leaves the judge to infer rules that may differ from the business policy.
Keep quality dimensions separate where they require different actions. A response with the correct policy but an unsuitable tone needs a different fix from one that quotes a nonexistent policy. Separate properties help identify that distinction and avoid a single composite number masking a serious failure.
Prompt properties support configuration of judges, categories and examples, and their definitions can be imported or exported as JSON. That makes the evaluation criteria an artifact the team can preserve alongside an application change. A prompt experiment is hard to assess if both the application and its scoring rubric change without a record.
Deepchecks can produce estimated Good or Bad annotations from property signals and configured rules. These labels are useful for sorting interactions and locating cases to inspect. They are estimates derived from the evaluation pipeline, so a Good label should not be treated as independent proof that an answer is correct.
Compare evaluator results with a small manually reviewed set. Include outputs that are correct but terse, plausible but false, partially correct and appropriately uncertain. If the judge rewards confident language more than supported content, refine the rubric or the data provided to it. The explanation accompanying a score can help reveal which interpretation produced the result.
Record disagreements rather than forcing every ambiguous case into a simple pass or fail. Some disagreements identify an unclear policy or missing source material. Others expose an evaluator weakness. Resolving the difference requires the application owner or a subject-matter reviewer, not only another automated score.
The agent evaluation guide describes Know Your Agent: configure an agent endpoint, build or generate scenarios, run simulations, capture execution traces and evaluate components. Supported integrations and SDK uploads can record agent, tool and model activity with their relationships.
This hierarchy helps distinguish an incorrect final answer from the step that caused it. A coordinator might choose a suitable tool, but a search component may retrieve irrelevant information. Alternatively, retrieval may succeed and a later model call may ignore the result. Inspecting individual spans makes those cases easier to separate than reading the final response alone.
The platform offers component and session views. Span-level properties can assess an individual action, while a session-level measurement considers the conversation as a whole. An agent can complete several sensible steps and still fail to deliver the user's requested outcome. Conversely, an ultimately useful answer may conceal a costly or unnecessary sequence of tool calls.
Use simulated scenarios with clear expected behavior and a controlled destination for actions. For a scheduling agent, test missing availability, conflicting constraints and a user changing their request. Keep any real external side effects outside the evaluation run unless the team explicitly intends and controls them. The scenario design determines what the simulation can establish.
The version comparison guide explains testing versions against the same inputs and using matching interaction identifiers for paired inspection. Teams can compare overall measurements and look at the same question's outputs side by side. Filters help locate differences in properties or estimated annotations.
Change one important variable at a time when diagnosing an improvement. If the prompt, model, retrieval index and evaluation rubric all change together, a better average score will not explain which change helped. Keep configuration notes and an identifiable dataset with each version so the result remains useful after the experiment.
Inspect regressions as well as improvements. An assistant might improve answers to common questions while becoming less reliable on exceptions. Review cases where the older version passed and the new one failed, and the reverse. A launch decision should consider the importance of those cases, not only the number of wins.
Maintain an evaluation set with routine requests, rare but consequential cases and examples from actual failures. Avoid continually tuning against the same handful of examples without adding new cases. A system can become better at a familiar test while leaving broader weaknesses unresolved.
Production monitoring applies the evaluation pipeline to production interactions and shows annotation distributions, property averages and trends. Teams can compare time windows and investigate changes. The documentation also describes forwarding evaluation signals to external observability tools such as Datadog and New Relic.
A change in a quality trend needs context. A new user population, updated documents or a different mixture of questions may affect the score even when the application code has not changed. Compare similar interaction types and check sample counts before attributing a movement to a specific release.
Cost tracking uses logged model names and token counts with configured model pricing. It supports interaction, session and version views, including aggregation through agent trace hierarchies. These calculations depend on correct logging and model-price configuration; an unmatched model name can leave a cost unavailable.
Use the cost view to identify the expensive path, then inspect its usefulness. A longer session may reflect productive work or repeated failed attempts. A low token cost can accompany a poor answer. Compare quality and operational measures together, and reconcile estimated model costs with provider billing when making a budget decision.
Consider an internal assistant that answers questions from company policy documents. The team wants to replace its prompt and verify that the change preserves correct handling of exceptions. Begin with a curated set of questions, the relevant approved sources and reviewer notes about the expected behavior.
A question whose answer is absent from the documents should have a defined expected behavior, such as requesting clarification or saying the source does not establish an answer. Otherwise an evaluator may reward a plausible invented response. The policy owner should define that standard before judging the candidate version.
This workflow can provide a repeatable record of a change and its tradeoffs. It does not make the evaluation dataset exhaustive. Keep a process for users to report errors and for reviewers to update the test set when a new policy or failure pattern appears.
Deepchecks currently offers paid Basic, Scale and Enterprise plans with a free trial. The pricing page distinguishes seats, applications, usage capacity, retention and support. It also describes dedicated and AWS-managed deployment options. Confirm the chosen deployment, procurement terms and feature access with the vendor.
SaaS usage is measured in Deepchecks Processing Units, or DPUs. The usage guide distinguishes data uploads and evaluation work and explains that this accounting applies to SaaS. Self-hosted and AWS arrangements have different model-cost responsibilities. A trial is an opportunity to evaluate the platform, not evidence of an ongoing free production plan.
Estimate usage using a representative trace and the properties the team actually needs. A long multi-agent session can contain many model calls and repeated context. Evaluating every field with every property may spend capacity on signals that do not inform a decision. Sampling and selectively pausing properties can help manage the workload.
Deployment choice also changes operational responsibilities. A managed service reduces infrastructure work; a private deployment requires appropriate configuration and ownership. Confirm data retention, model connections and access controls for the selected arrangement, especially when traces contain customer information or confidential documents.
Deepchecks fits teams comparing AI application versions, maintaining quality rubrics or diagnosing production agent behavior. AI engineers can investigate traces, quality teams can maintain evaluation standards, and product owners can review whether the system fulfills a business task. It is less immediately useful when there is no application, dataset or agreed quality definition to assess.
Automated judges can misunderstand domain-specific rules, and their results depend on the context supplied. Missing retrieval documents can make a grounding judgment uninformative. Missing spans can conceal a failing component. Good integration and calibration are prerequisites for interpreting a dashboard sensibly.
Scores and labels organize evidence; they cannot replace required business approvals or exhaustive testing. Some failures require deterministic checks, such as verifying an identifier format or enforcing permission boundaries. Use those alongside qualitative evaluation rather than expecting an LLM judge to establish every property of a system.
Its principal role here is evaluation, observability and monitoring of an application or agent. The team connects the system and supplies or generates test scenarios for assessment.
Yes. Log each version against shared inputs and matching identifiers for detailed paired comparison. Keep evaluation settings consistent enough to interpret the differences.
No. They are estimated annotations based on property signals and rules. Calibrate them with reviewed cases and inspect consequential disagreements.
The current commercial platform offers a trial and paid plans. Deepchecks also has open-source testing resources, which should be distinguished from the hosted LLM Evaluation subscription.