Can your measurement system support the claim you are making?
Thirdplane produces dated, evidence-limited assessments of the judges, verifiers, reward signals, labels, and selection policies behind consequential AI claims.
Request a research conversation
Teams selling RL environments, training data, or evaluation systems that another organization must rely on.
The claim at stake
Scores train, rank, and sell AI systems. But once a model or vendor optimizes against a score, it can stop measuring the outcome a customer relies on.
When a customer, procurement team, or counterparty must trust a claim; when a release, training, purchase, or public decision depends on it; and when ordinary tests cannot fully settle the human or semantic outcome at stake.
If a claim is fully and objectively checkable, such as whether a date was extracted correctly or a test passed, you may not need us.
What Thirdplane can establish
A dated, evidence-limited assessment of whether a specified measurement system supports the claim being made, including where it does not.
The evidence path changes
The claim rests on the system owner’s evidence.
- The owner defines the task, evidence, and success rule.
- The boundary of what the score supports is left to inference.
- A buyer has no separate path to inspect scope or limits.
A specified claim has an inspectable evidence path.
- A dated record identifies the version and evidence reviewed.
- Tests and a separate reference path are defined for one claim.
- Findings state both the evidence of support and the limits.
What we examine
Judges, verifiers, reward signals, label sets, and selection policies are different parts of the same measurement system. The relevant question is whether that specified system supports a specified claim for the population and decision at hand.
How we work
- Freeze the version.
- Define the test.
- Compare against an independent reference.
- State the limits.
Why this work
Thirdplane Labs is led by Jasmine Poon, who spent nearly five years at Ontra across AI/ML product, platform, and infrastructure.
She built retrieval, evaluation, and LLM-as-a-judge systems for Ontra’s AI organization. That experience made the measurement problem practical.
Research in progress
Who checks the judge? The missing layer in AI evaluation.
Why scores need evidence beyond self-attestation when people optimize against them.
We are assessing whether public Terminal Bench artifacts can support a dated, reproducible study. It would not certify the benchmark.
Tell us what the score is meant to measure.
We are conducting a small number of 20-minute research conversations with teams whose customers or counterparties rely on an evaluation claim. This is not an automated sales funnel or a commitment to an engagement.