thirdplane labs

AI evaluation services

Contact

Does your AI evaluation measure what matters?

Thirdplane Labs provides evaluation design and review for teams building AI products.

I investigate where AI and human judgments diverge, so your team can decide what to trust, what to improve, and what needs review.

Where I can help

  • Your scores look good, but important failures still get through.

    An assessment of what the evaluation misses, with supporting examples and priorities for correction.

  • Your team can’t reliably tell whether a change is better.

    A working evaluation with relevant cases and criteria for comparing models, prompts, or workflows.

  • Manual review is slowing development.

    A workflow that automates suitable checks and routes uncertain cases, with supporting evidence, to human review.

How an engagement works

We start with one decision your team needs to make, then agree on the scope, relevant examples, and a useful deliverable. I work with your existing product and evaluation setup, and hand over the agreed work with findings, limitations, and next steps.

Different disagreements call for different checks

The value is in finding what needs to change, not simply making the scores agree.

  • AI–AI

    Can you rely on the evaluator?

    Different AI judges -- or repeated runs of the same judge -- may score the same work differently. Examine variability and bias: is the verdict sensitive to the model, prompt, or answer order?

    Identify unstable judgments and test the evaluator against task evidence.

  • AI–human

    Does the evaluator reflect what matters?

    An AI verdict may differ from a human review because of missing context, different criteria, or an error on either side. Inspect the mismatches before treating either judgment as the reference.

    Calibrate review with concrete examples and define which cases need human judgment.

Human–human

People can disagree too. Clarify criteria and investigate disputed evidence; preserve defensible differences where the task allows them. Human review needs a clear standard of its own.

Consistency needs evidence

Any of these disagreements can reveal where to investigate. For each, ask: are judgments consistent under comparable conditions, and does independent evidence connect them to the intended outcome?

  • Calibrate Stronger evidence · Less consistent judgments

    Useful evidence, uneven judgments.

    Judgments vary despite useful outcome evidence.

    Check how reviewers interpret the evidence and criteria.

  • Target state Stronger evidence · More consistent judgments

    Dependable evaluation.

    Consistent judgments, supported by independent outcome evidence.

    Revalidate as the system changes.

  • Investigate Weaker evidence · Less consistent judgments

    Unclear signal.

    Judgments vary, and their link to the intended outcome is unclear.

    Review the criteria, evidence, and scoring process.

  • Validate Weaker evidence · More consistent judgments

    Consistent, but unvalidated.

    Judgments agree; their connection to the intended outcome is still uncertain.

    Test against independent outcome evidence.

A conceptual framework for the task and cases examined, not measured results. Agreement alone does not establish correctness; some differences in judgment are legitimate.

Open the full-size diagram

Sources behind the framework

The framework is Thirdplane Labs’ synthesis, informed by Zheng et al. on model-judge biases, Jiang and de Marneffe on human disagreement, and Anthropic on complementary evaluation methods.

What decision are you working toward?

Tell me what you’re building, what you need to decide, and where your current evaluation falls short.

Email j@thirdplane.io