Biopharma Bench:
Can Frontier Agents Do Real Professional Work?

by Raycaster in Research
Company environments
12
Private biopharma records
Professional assignments
71
Employee seats, not prompts
Frozen criteria
752
15+ year director standard
Best full-task Pass@1
2.8%
Grok 4.6 · 2 of 71
Biopharma Bench V0.1 Suite· 71 tasks across 12 companies · 752 criteria
Share

Frontier models can draft convincing technical prose in minutes. The harder test is whether an agent can navigate contradictory company records, determine which specifications govern, and complete an assignment without introducing critical defects. Today we are releasing Biopharma Bench V0.1: 71 professional assignments across 12 biopharma company environments.

AI agents have become remarkably good at well-specified tasks. When given a clear objective, a curated document, and explicit instructions, frontier models can draft fluent, technically persuasive prose in minutes.

But real professional work rarely arrives in a clean prompt. In drug and medical device development, the hardest part of the job is often figuring out what the ground truth actually is. Evidence is scattered across dozens of file directories, contradictory batch records, superseded protocols, and informal team messages. An exploratory lab memo might suggest a promising hypothesis while an approved specification legally prohibits it. A supplier might assert that a replacement film is identical to an approved component, but their own test data was measured under the wrong storage conditions.

Doing the job requires investigating those discrepancies, determining which document actually carries regulatory authority, and making decisions that can survive scrutiny from an inspector.

That is the question behind Biopharma Bench V0.1: can frontier agents handle the messy, end-to-end reality of professional biopharma work?

We placed frontier models into employee seats across 12 biopharma company environments built from genuine product histories and private industry records. Rather than answering isolated questions, each agent sits down at an employee's desk with a full file room, controlled document registers, and colleagues on chat. We evaluated each system across 71 professional assignments against 752 frozen, expert-authored substantive criteria.

The headline finding: frontier agents still have a long way to go before operating independently in regulated industry. Across 71 scored assignments, complete task success was rare: systems achieved between 0% and 2.8% Pass@1 (satisfying every substantive criterion on the first trial). On average rubric coverage (macro task mean), GPT-6 Astra led with 65.9%, followed by Claude Opus 5 at 63.8% and Grok 4.6 at 63.0% (which took the top spot for complete deliverables with 2 full passes).

Pass@1 measures complete work: did the agent satisfy every frozen substantive criterion without introducing a single critical defect? Macro task mean measures average rubric coverage. Neither represents regulatory sign-off, and neither means an AI system is ready to operate without accountable human oversight.

The limits of clean-prompt benchmarks

Imagine if an AI agent had to go back in time, sit in the regulatory and CMC seats of an innovative drug company, navigate years of conflicting internal records and agency requests, and bring a breakthrough therapy—like a first-in-class GLP-1 receptor agonist or an antibody-drug conjugate—through rounds of regulatory approval. Could it autonomously do the job without making a fatal mistake?

Most existing agent benchmarks evaluate models in simplified, synthetic conditions: a short prompt, a pre-extracted text snippet, and a narrow set of expected answers. While useful for measuring basic reasoning or coding syntax, these setups miss the core friction of professional knowledge work in regulated industry. What makes a benchmark valid? We anchor Biopharma Bench on three uncompromising pillars:

1. World Realism (Messy evidence over pre-packaged context). Real work does not arrive in a single prompt folder. In biopharma development, evidence is distributed across hundreds of files: regulatory filings, quality deviation logs, clinical trial protocols, LIMS assay exports, and informal chat correspondence. An agent cannot simply summarize what is in front of it; it must search file trees, recognize which specifications are legally obsolete, and piece together the factual timeline itself amidst realistic operational noise.

2. Task Realism (Controlling authority over surface fluency). Language models are trained to produce confident, coherent prose. In regulated development, plausible writing is dangerous if it rests on the wrong authority. An informal team memo cannot override an approved Common Technical Document (CTD) specification, and an exploratory Phase 1 protocol does not govern commercial release testing. Tasks require producing genuine, end-to-end deliverables—reconciling analytical deviations, drafting responses to binding agency clinical holds, or justifying commercial shelf life—where determining which document actually governs is the primary duty of an employee in that seat.

3. Rubric Realism & Ground Truth (Evaluated against statutory truth, not past human error). A great benchmark does not evaluate by fuzzy cosine similarity or naive string matching against historical human drafts. Real human teams operated under intense deadlines with imperfect information, occasionally making arithmetic errors, missing safety warnings, or taking regulatory shortcuts. Biopharma Bench rubrics evaluate deliverables against governing regulatory mandates (US FDA 21 CFR, ICH guidelines, pharmacopeial standards) and immutable scientific laws. If an agent catches a stoichiometric error, identifies an unaddressed safety signal, or applies stricter statistical rigor than the original company did, the rubric rewards it. We evaluate against regulatory and scientific truth, never penalizing an AI system for exceeding historical human performance.

When a benchmark tests an agent on a curated prompt, it tests transcription and summarization. What makes an evaluation valid is whether an agent can uncover the facts, resolve contradictory data, and defend its conclusions against real statutory and operational standards.

Benchmark contents & company environments

The benchmark roster contains 71 assignments across 12 distinct company environments, built from authentic development milestones and private biopharma records. These environments span major health authorities (US FDA CDER, CBER, CDRH; the European Medicines Agency; Health Canada; Korea MFDS; and China NMPA / CDE), diverse modalities (mRNA vaccines, antibody-drug conjugates, targeted small molecules, peptide injectables, oral solid fixed-dose combinations, and IVD companion diagnostics), and essential company functions across regulatory affairs, quality compliance, MSAT, and clinical site operations.

Rather than answering an abstract question or parsing an isolated snippet, the agent sits down at an employee's desk at a specific historical moment in the company's journey. It receives the actual files, historical revisions, and background noise that a person in that seat would have had. Its job is to investigate the records and produce the required work product: an updated regulatory assessment, an out-of-specification investigation, or an audit sign-off package.

Alongside these private environments stands Havenor Therapeutics, a completely synthetic showcase company created so researchers can openly inspect full trajectories, coworker interactions, and deliverables in public without exposing confidential partner files. Havenor's calendar runs to September 2026 without faketime. Havenor serves as an open testbed and is evaluated separately from the sealed 71-task benchmark.

You can explore curated run trajectories and actual deliverables from Havenor on the live benchmark page. The public Havenor desks and rubrics are on Hugging Face. To run the tasks with records, LIMS, and the verifier, use Harbor.

Leaderboard & results

Figure 1 reports the headline leaderboard across all eight frontier systems. Rank is determined bymacro task mean—the primary metric, where every assignment carries equal weight. We also report Pass@1 (share of assignments where every single frozen substantive criterion passed on the first trial), absolute criteria passed, mean token cost, ATIF steps, tool calls, and execution time. You can sort the leaderboard by clicking any column header.

Figure 1 · Benchmark Leaderboard

Benchmark results across eight frontier models

8 systems · 71 tasks · 12 companies · 752 criteria · 1 scored trial per task

Click a column to sort
Biopharma Bench V0.1 scores across eight frontier models.
ModelAgentThinking
1GPT-6 AstraCodexmedium
65.9%
0 / 71
499/752$3.3226.320.36.5 min
2Claude Opus 5Claude Codemedium
63.8%
0 / 71
485/752$2.6829.229.410.6 min
3Grok 4.6Cursor CLIhigh
63.0%
2 / 712.8%
480/752$1.1620.546.78.8 min
4DeepSeek V4.1 FlashPihigh
57.0%
0 / 71
433/752$0.1136.240.56.3 min
5Kimi K3Cursor CLI / Pidefault
48.5%
1 / 711.4%
374/752$0.7221.925.14.9 min
6GLM-5.3 FlashPihigh
47.6%
0 / 71
365/752$0.0422.825.83.1 min
7Gemini 3.8 FlashCursor CLIhigh
46.6%
0 / 71
356/752$1.1275.277.214.5 min
8GPT-5.6 SolCodexmedium
46.6%
0 / 71
364/752$1.2729.223.24.1 min
Results are subject to variance; one scored trial per task. Macro mean gives equal weight to each assignment. Pass@1 is the percentage of assignments where every frozen substantive criterion passed on the first attempt. Steps, tool calls, and time are the average across all 71 tasks, counted from the scored trial logs. Time is agent execution time. Candidate cost reflects measured token usage priced using public provider API rates; grader execution spend is excluded.

Across 71 scored trials per system, GPT-6 Astra tops the leaderboard with a 65.9% macro task mean (passing 499 of 752 criteria), followed closely by Claude Opus 5 at 63.8% (485 criteria) and Grok 4.6 at 63.0% (480 criteria). Grok 4.6 achieved the highest rate of complete work, recording 2 full passes (2.8% Pass@1, successfully clearing both an NMPA methotrexate deficiency response and a Korea ADC manufacturer review). DeepSeek V4.1 Flash reached 57.0% across all 71 tasks, Kimi K3 scored 48.5% with 1 full pass, GLM-5.3 Flash achieved 47.6%, and Gemini 3.8 Flash tied GPT-5.6 Sol at 46.6%(356 and 364 criteria passed, respectively).

Figure 2 contrasts this average rubric coverage with complete, unassisted success (Pass@1). While models regularly earn partial credit by drafting passable introductory sections or reciting background literature, clearing every single substantive check without introducing a critical compliance defect proved exceptionally difficult across all providers.

Figure 2 · Partial vs. Full Success

Partial rubric coverage versus full-task completion

8 systems · 71 tasks · 12 companies · 752 criteria

Sorted by macro mean (highest first)
Pass@1 (100% criteria)Macro task mean

Sorted by Macro mean, highest first.

01GPT-6 Astra

71/71 tasks · Codex · medium thinking
Pass@1
0.0%
0/71 full passes
Macro mean
65.9%
Equal weight per task
Criteria
499/752
66.4% of items

$3.32 per task · 26.3 steps on 71 tasks · 20.3 tool calls · 6.5 min

02Claude Opus 5

71/71 tasks · Claude Code · medium thinking
Pass@1
0.0%
0/71 full passes
Macro mean
63.8%
Equal weight per task
Criteria
485/752
64.5% of items

$2.68 per task · 29.2 steps on 71 tasks · 29.4 tool calls · 10.6 min

03Grok 4.6

71/71 tasks · Cursor CLI · high thinking
Pass@1
2.8%
2/71 full passes
Macro mean
63.0%
Equal weight per task
Criteria
480/752
63.8% of items

$1.16 per task · 20.5 steps on 71 tasks · 46.7 tool calls · 8.8 min

04DeepSeek V4.1 Flash

71/71 tasks · Pi · high thinking
Pass@1
0.0%
0/71 full passes
Macro mean
57.0%
Equal weight per task
Criteria
433/752
57.6% of items

$0.11 per task · 36.2 steps on 71 tasks · 40.5 tool calls · 6.3 min

05Kimi K3

71/71 tasks · Cursor CLI / Pi · default thinking
Pass@1
1.4%
1/71 full passes
Macro mean
48.5%
Equal weight per task
Criteria
374/752
49.7% of items

$0.72 per task · 21.9 steps on 71 tasks · 25.1 tool calls · 4.9 min

06GLM-5.3 Flash

71/71 tasks · Pi · high thinking
Pass@1
0.0%
0/71 full passes
Macro mean
47.6%
Equal weight per task
Criteria
365/752
48.5% of items

$0.04 per task · 22.8 steps on 71 tasks · 25.8 tool calls · 3.1 min

07GPT-5.6 Sol

71/71 tasks · Codex · medium thinking
Pass@1
0.0%
0/71 full passes
Macro mean
46.6%
Equal weight per task
Criteria
364/752
48.4% of items

$1.27 per task · 29.2 steps on 71 tasks · 23.2 tool calls · 4.1 min

08Gemini 3.8 Flash

71/71 tasks · Cursor CLI · high thinking
Pass@1
0.0%
0/71 full passes
Macro mean
46.6%
Equal weight per task
Criteria
356/752
47.3% of items

$1.12 per task · 75.2 steps on 71 tasks · 77.2 tool calls · 14.5 min

Sort by macro task mean, full passes, or criteria passed. Macro mean gives every assignment equal weight. Pass@1 is the share of assignments where every frozen substantive criterion passed in the single scored trial. Criteria passed is the item count across all 752 checks. Black marks Pass@1. Blue marks the macro mean. Cost is the mean candidate spend on all 71 tasks. Steps, tool calls, and time are the average across all 71 tasks, counted from the scored trial logs. Time is agent execution time.

Because Biopharma Bench runs one trial per task, gaps of a few points (such as Claude Opus 5 vs. Grok 4.6, or Gemini 3.8 Flash vs. GPT-5.6 Sol) may not be statistically meaningful. Repeat trials are planned; in our experience across benchmark tasks, variance is typically within ±3–5% after three trials.

However, looking only at top-line averages conceals how agents actually perform on the job. Because each environment represents an authentic professional seat, models are stronger on different jobs:

Figure 3 · Domain Comparisons

Complementary capabilities across specialized biopharma seats

Performance variations across regulatory and technical domains

mRNA vaccine lifecycle (FDA CBER)

Astra vs. Opus lead

Outside the CBER mRNA vaccine environment, Astra and Opus pass the exact same number of items (445 of 685). Astra's 14-item overall lead comes entirely from the CBER mRNA vaccine environment (54/67 vs. 40/67).

CDRH IVD & Synthetic API

Opus vs. Grok comparison

Opus leads the FDA CDRH companion diagnostic (44/67 vs. 31/67) and the Health Canada peptide injectable freeze (24/32 vs. 18/32), while Grok leads the FDA synthetic API inspection packet (52/60 vs. 41/60). Overall, Opus leads Grok by just 5 items.

Korea MFDS (ADC)

Three-way parity

Korea MFDS is a virtual tie across three architectures: Opus (54/78), Grok (53/78), and DeepSeek (52/78), with Astra scoring 44/78.

Site Notes & Metformin

Gemini vs. Sol tie at 46.6%

Gemini and Sol tie on the macro mean at 46.6%, but solve opposite seats: Gemini leads on clinical site notes (92/127 vs. 72/127) and a China NMPA metformin FDC reply (30/48 vs. 16/48), while Sol leads on Korea ADC, CDER/EMA targeted small molecules, CBER mRNA vaccines, and CDRH companion diagnostics.

Models display distinct domain profiles rather than uniform advantage across all seats. Systems with similar overall macro means frequently excel at different regulatory or clinical tasks.

Take Astra and Opus: outside the mRNA vaccine lifecycle at CBER, the two models pass the exact same number of criteria (445 of 685). Astra's entire 14-item lead comes from the CBER mRNA vaccine environment (54/67 vs. 40/67). Opus leads the FDA CDRH companion diagnostic environment (44/67 vs. Grok's 31/67), while Grok pulls ahead on the FDA synthetic API inspection response (52/60 vs. Opus's 41/60). On Korea MFDS ADC review, Opus (54), Grok (53), and DeepSeek (52) finish in a virtual dead heat. And while Gemini 3.8 Flash and GPT-5.6 Sol tie on the overall macro mean at 46.6%, they solve completely opposite seats: Gemini leads on clinical site notes (92 vs. 72) and a China NMPA metformin FDC reply (30 vs. 16), whereas Sol leads on Korea ADC, CDER/EMA targeted small molecules, CBER mRNA vaccines, and CDRH companion diagnostics.

Scores only tell half the story. Figure 4 plots mean score against mean candidate cost for all 71 tasks. The dashed line marks models where no other model is both cheaper and higher scoring: GLM-5.3 Flash at $0.04 (47.6%), DeepSeek V4.1 Flash at $0.11 (57.0%), Grok 4.6 at $1.16 (63.0%), Claude Opus 5 at $2.68 (63.8%), and GPT-6 Astra at $3.32 (65.9%). Kimi, Gemini, and Sol cost more than a higher-scoring model. Candidate cost is measured token use priced at published API rates. Grading is excluded.

Figure 4 · Cost-Performance Frontier

Cost-performance Pareto frontier across the 71-task benchmark

45%50%55%60%65%70%$0.00$1.00$2.00$3.00$4.00Mean task score ↑Mean candidate cost per task (USD) →GLM-5.3 FlashPi · 47.6% · $0.04DeepSeek V4.1 FlashPi · 57.0% · $0.11Grok 4.6Cursor CLI · 63.0% · $1.16Claude Opus 5Claude Code · 63.8% · $2.68GPT-6 AstraCodex · 65.9% · $3.32
GLM-5.3 Flash ★
Pi47.6% · $0.04
DeepSeek V4.1 Flash ★
Pi57.0% · $0.11
Kimi K3
Cursor CLI / Pi48.5% · $0.72
Gemini 3.8 Flash
Cursor CLI46.6% · $1.12
Grok 4.6 ★
Cursor CLI63.0% · $1.16
GPT-5.6 Sol
Codex46.6% · $1.27
Claude Opus 5 ★
Claude Code63.8% · $2.68
GPT-6 Astra ★
Codex65.9% · $3.32
Cost-performance frontier across all 71 benchmark tasks. The dashed line connects systems on the non-dominated cost-performance frontier (GLM-5.3 Flash, DeepSeek V4.1 Flash, Grok 4.6, Claude Opus 5, and GPT-6 Astra). Hollow points are dominated. The score axis starts at 45%, not 0%. Costs reflect candidate model execution token usage priced using published rates.

Real-world failure modes & traps

When an agent falls short on these tasks, it is almost never because the writing looks unpolished. The prose is usually fluent, confident, and neatly formatted. The failure almost always comes down to authority: which document actually governs?

You can see this clearly in Havenor Therapeutics Task 03, where Claude Opus 5 (running in Claude Code) was assigned as a principal engineer to review sterile-filter validation studies and draft a technical brief ahead of a health authority audit.

Figure 5 · Havenor Therapeutics Task 03

Three records. One claim that does not survive the protocol.

  1. 01 · ReportAVT-VAL-107

    Bacterial retention report

    Within the protocol-defined 30-minute window; organism viability control conformed.

    Observation VAL-FLT-BRT-002-OBS-01. The second replicate began 14 minutes after the nominal inoculation window. The report records it as an accepted observation.

  2. 02 · Governing protocolAVT-VAL-107-P

    Approved protocol v01

    The approved protocol contains no 30-minute inoculation window.

    Its acceptance criteria cover organism identity, challenge density, filtrate recovery, recovery controls, and a bubble-point limit. None of those criteria define a 30-minute inoculation window.

  3. 03 · Candidate briefDefensibility review

    Inspector-facing brief

    The brief repeats the 30-minute window as protocol-defined.

    In the published Havenor Task 03 trial, the candidate opened the retention report and the scanned protocol, then carried the report’s claim into the final review. The brief does not record that AVT-VAL-107-P contains no such window.

Three records from the Havenor Task 03 showcase. In observation VAL-FLT-BRT-002-OBS-01, the secondary report claims an inoculation delay was “within the protocol-defined 30-minute window.” But the approved protocol (AVT-VAL-107-P) defines no such requirement. The agent opened both documents, missed the discrepancy, and repeated the unsupported claim straight into the final brief for the inspector.

The model demonstrated real capability: it navigated complex folder hierarchies, opened validation reports, and extracted data from executed protocol scans that simple text extractors couldn't parse.

The breakdown was subtler—and far more dangerous. In a secondary retention report, an observation noted that an inoculation began 14 minutes late, claiming this was “within the protocol-defined 30-minute window.” Yet when you open the controlling protocol (AVT-VAL-107-P), no such 30-minute window exists anywhere in the document.

An experienced human engineer knows that a secondary report cannot invent a protocol requirement out of thin air. If it isn't in the protocol, you flag the discrepancy. Instead, the model took the secondary report at its word, repeated the phantom 30-minute window as an established fact, and carried it straight into the brief prepared for the inspector.

In a real audit, handing an inspector a brief that misstates or invents a protocol requirement creates immediate compliance exposure. This illustrates the central finding of the benchmark: an agent can search thoroughly, reason through complex data, and write beautifully—and still fail at the essential act of professional verification.

Case Studies · Ground Truth vs. Plausible Fiction

Five industry failure modes: Where plausible prose creates audit failure

The Binding Hold vs. The Internal Steering Memo

US FDA CDER · IND Hold Response
What the agent did (Plausible prose)

When tasked with responding to the hold, multiple frontier agents followed an internal executive steering memo that urged 'preparing trial sites quietly in parallel' while drafting the rebuttal. The models drafted a plan to initiate clinical site setup and screening immediately.

Real-world consequence (Regulatory failure)

Under 21 CFR 312.42, a clinical hold is legally binding upon receipt. Enrolling or even screening healthy volunteers before written FDA removal constitutes a federal violation that jeopardizes the entire development program and subjects the sponsor to formal regulatory enforcement.

Evaluation principle: Ground truth is not an internal memo or executive enthusiasm. An agent must distinguish between what a team wants to do and what governing regulatory statute legally permits.
Real-world failure cases reconstructed from actual regulatory filings and audit findings. In each scenario, models produced well-formatted, confident responses that would fail an expert review or trigger formal regulatory action.

In the history of biopharma, catastrophic setbacks are rarely caused by failures of basic prose; they are caused by subtle disconnects between physical reality, vendor operations, and statutory regulations. When Genzyme faced a consent decree and a $175M fine at its Allston Landing facility in 2009, the viral contamination was rooted in sanitary piping dead-legs and missed environmental warnings that previous inspection responses had glossed over. When Abbott had to abruptly withdraw Norvir (ritonavir) capsules in 1998, it was because an unforeseen, thermodynamically stable crystalline polymorph nucleated in the manufacturing line, cutting solubility in half. And when commercial launches like high-dose Eylea faced sudden Complete Response Letters in 2023, the blocker was not clinical efficacy, but unaddressed inspection findings at a third-party contract fill-finish facility.

In every case, the critical facts existed in internal records or vendor audit trails before the crisis broke. A vigilant professional caught the discrepancy early; a hasty team overlooked it. That is why Biopharma Bench evaluates against statutory and scientific ground truth rather than historical human consensus. In retrospect, even experienced teams working under extreme deadlines made sub-optimal assumptions or missed contradictory data. Our rubrics never penalize an AI agent for demonstrating superior scientific vigilance, correcting historical stoichiometric errors, or refusing to pool non-homogeneous stability batches—even if the historical sponsor originally took the shortcut.

How we evaluated

Every benchmark attempt takes place inside an isolated container. The agent is assigned an employee seat with access to the local company files, relevant tools, and a standard set of instructions: explore the company records, make reasonable inferences when encountering ambiguity, and continue working unless genuinely blocked.

We evaluate complete model-and-harness pairings, not bare models in a vacuum. On the 71-task roster, Astra and Sol ran on Codex at medium thinking, Opus on Claude Code at medium thinking, Grok and Gemini on Cursor CLI at high thinking, DeepSeek and GLM on Pi at high thinking, and Kimi on a blend of Cursor CLI and Pi at the provider default. DeepSeek has no medium setting, and the Grok and Gemini seats are the high model slugs. The harness governs how an agent inspects files, runs tools, and manages context, so results reflect the combined system.

Figure 6 · Evaluation Flow

How an assignment becomes a benchmark score

01 · Staged assignment

Employee seat & records

Role-appropriate seat, realistic company desk, files, and authorized tools.

02 · Candidate run

Model + harness

Autonomous investigation under standard non-clarification instructions.

03 · Work product

Deliverable

Document, spreadsheet, or authorized system action.

Hidden grading boundary · Frozen criteria applied only after the runQuarantined from candidate
Task score (share of criteria passed)

Mean task score

Share of frozen substantive criteria passed for that task, macro-averaged equally across all assignments in the benchmark suite.

All-criteria outcome (every criterion passed)

Pass@1

The percentage of assignments where the candidate satisfied every frozen requirement in its single scored trial without human intervention.

Flow from staged assignment to post-run grading. Deliverables are evaluated against frozen criteria that remain hidden from the candidate during execution. Substantive criteria are scored by an automated rubric grader (GPT-5.6 Luna, medium thinking) in Harbor, alongside deterministic checks for narrow quantitative claims.

Once an agent finishes, its deliverable is scored against frozen criteria that were quarantined away during the run. When a check is strictly quantitative—such as a specific calculated yield or a key spreadsheet value—we evaluate it with deterministic code.

Most professional work, however, cannot be captured by a regex. Two senior regulatory writers might structure a justification memo differently while both making sound, defensible arguments. Deliverables are scored using an automated rubric grader powered by GPT-5.6 Luna (medium thinking) executed within the Harbor evaluation harness. The grader evaluates each substantive criterion and aspect check against the submitted work alongside the underlying source documents and frozen truth specifications. A check passes only when the submitted work itself demonstrably supports it.

A brief word on what these scores mean: strong performance shows that an agent can navigate complex internal records, reconcile conflicting evidence, and draft defensible technical work. It does not certify regulatory compliance, and it does not mean an agent is ready to operate without accountable human oversight.

What comes next

This release represents an initial snapshot across eight leading systems. As frontier models and agent harnesses continue to evolve, Raycaster will publish updated results and expand the benchmark roster.

We believe evaluations are only as credible as their transparency. Full run details, candidate transcripts, and failure analysis across all environments are available upon request. For Havenor Therapeutics, we invite you to inspect the full transcripts, coworker interactions, and final deliverables for yourself on the live benchmark page.

Inspect the work, not just the score

Explore the Havenor trials, step-by-step agent trajectories, and complete deliverables for Biopharma Bench V0.1 on Raycaster Eval.