TRACES is Apodex’s benchmark for open ended scientific discovery: instead of checking a response against a fixed answer key, it evaluates how an AI system investigates, uses tools, responds to feedback, and supports i... The benchmark evaluates six capabilities—Tools, Repair, Alternatives, Coherence, Evidence, and S...
Research answer

Create a landscape editorial hero image for this Studio Global article: What is Apodex’s TRACES benchmark, launched to evaluate “discoverative AI,” how does it move beyond traditional answer-key benchmarks by pla. Article summary: TRACES is Apodex’s “reality benchmark” for evaluating discoverative AI: systems intended to investigate open scientific problems and reach potentially unknown conclusions, rather than retrieve a predefined answer. It eva. Topic tags: general, academic, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, char
AI benchmarks have traditionally asked systems to produce an answer that can be compared with a known key. Apodex’s TRACES takes a different approach: it places a complete AI “heavy-duty solver” inside an executable, stateful environment and observes how it investigates an open scientific problem. The system may use tools, receive feedback, revise failed work, and pursue conclusions that are not already settled.
The distinction matters because a plausible final answer can be lucky, poorly supported, or impossible to reproduce. TRACES is designed to evaluate the investigation behind the result as well as the result itself.
TRACES is part of Apodex Discovery, a framework for building and evaluating what Apodex calls “discoverative AI”: systems aimed at reaching potentially new conclusions rather than simply retrieving or restating known information. The benchmark uses executable environments, episodes, tasks, and hidden verifiers to make open-ended investigations testable.
A submission is not just a foundation model. The evaluated unit is a complete solver, including the model, orchestration harness, tools, memory, and control policies used to conduct the investigation.
Conventional benchmark tasks generally define the question, available inputs, and expected answer in advance. TRACES instead starts with a real-world ambition that must be translated into a tractable investigation. The solver interacts with an environment, takes actions, calls scientific or computational tools, and receives experiment-like or verifier feedback.
That creates a longer evaluation loop:
The benchmark therefore records the trajectory of the investigation, not merely whether the final output resembles a reference answer. This is especially important when the underlying scientific question has no settled answer that can serve as a conventional ground truth.
The acronym names six process-verification dimensions, sometimes described as HDS6:
Together, these dimensions are intended to distinguish a reliable investigation from a confident but unsupported guess. They also allow reviewers to assess the quality of a solver’s process when a definitive answer is not yet available.
Apodex’s Discovery framework identifies several requirements for evaluating open-ended discovery: a well-formed problem, a reality-based environment, mechanisms for verification, and loops that allow the system to repair its work. TRACES is the benchmark implementation of that approach.
For its problem-scouting process, Apodex says it surveyed 561 industries across 16 sectors and assembled 423 high-value real-world problems. The initial release selected 20 problems for executable evaluation.
The available materials identify examples spanning:
These examples show the benchmark’s intended breadth, but the provided materials do not establish a complete public inventory of every task or domain in the release.
AI systems are increasingly being used—or proposed—for tasks such as generating hypotheses, selecting research directions, analyzing experimental results, and helping design iterative workflows. In these settings, fluency is not enough. A system must know what to test, interpret what happened, recover from mistakes, and avoid claiming more than its evidence supports.
TRACES addresses that gap by treating scientific investigation as an interactive process. A solver that reaches a useful result through traceable evidence and appropriate revisions is meaningfully different from one that produces a similar conclusion without a sound path to it.
The approach does not eliminate human oversight. Researchers and domain experts still need to set objectives and constraints, assess ethical and practical risks, validate assumptions and evidence, and decide whether a conclusion justifies real-world action. TRACES’s emphasis on inspectable trajectories and bounded claims is valuable precisely because it gives reviewers more to examine than an isolated final assertion.
Apodex’s public announcement describes an open call for both solver systems and scientific problems.
A solver submission would represent the full system used to investigate a task—not only the underlying model. That includes the harness or orchestration layer, tools, memory, and control policy.
A problem proposal would need to be expressible as an executable and verifiable setting. In practice, that means defining what the solver can observe and do, what feedback it receives, and how intermediate or final claims can be checked against explicit success criteria.
TRACES changes the benchmark question from “Did the AI give the expected answer?” to “Did the AI conduct a disciplined, evidence-grounded investigation?” That makes it a closer fit for scientific discovery than static answer-key tests, while also exposing a central limitation: process scores can improve evaluation, but they cannot by themselves establish that an AI-generated discovery is safe, useful, or ready for deployment. Human scientific judgment remains essential.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
TRACES is Apodex’s benchmark for open ended scientific discovery: instead of checking a response against a fixed answer key, it evaluates how an AI system investigates, uses tools, responds to feedback, and supports i...
TRACES is Apodex’s benchmark for open ended scientific discovery: instead of checking a response against a fixed answer key, it evaluates how an AI system investigates, uses tools, responds to feedback, and supports i... The benchmark evaluates six capabilities—Tools, Repair, Alternatives, Coherence, Evidence, and Scope—so a system can be judged even when no definitive ground truth answer is available.
Teams can participate by submitting a complete solver system or proposing a scientific problem that can be converted into an executable, verifiable environment.