OpenAI reported on June 17, 2026, that GPT-Rosalind passed 36.1 percent of the 750 tasks in its new LifeSciBench evaluation. The result is evidence of limited performance on this benchmark, not a finding that the model completed 36.1 percent of real laboratory work or reached 36.1 percent of a scientist's ability.

OpenAI says LifeSciBench tests seven research workflows across seven biological domains using tasks written by 173 scientists with doctoral training and biotechnology or pharmaceutical industry experience. The same organization designed the benchmark, evaluated three of its own models among the five systems and published the results, so the scores require both technical and institutional scrutiny.

An AI system is evaluated on life-science research tasks in a laboratory setting (illustrative image)

The pass rate is a rubric threshold, not simple accuracy

LifeSciBench does not use multiple-choice questions or a single reference answer. Each task combines a free-response prompt with relevant context or artifacts and a task-specific rubric. Its 750 tasks contain 19,020 rubric criteria, an average of 25 per task, covering factual claims, calculations, decisions, justification, caveats and formatting.

The LifeSciBench paper defines a pass as a response receiving at least 70 percent of the available rubric points for a task. A model can earn partial credit without passing. Calling 36.1 percent an ordinary accuracy score would therefore obscure both the scoring rule and the partial work captured by the separate normalized rubric score, on which GPT-Rosalind averaged 0.576.

Sequencing equipment and laboratory instruments used in life-science research (illustrative image)

Files and figures produced the clearest performance drop

The overall ranking put GPT-Rosalind first with a 36.1 percent pass rate. GPT-5.5 reached 25.7 percent, Gemini 3.1 Pro 23.6 percent, GPT-5.4 20.7 percent and Grok 4.3 13.0 percent. None of the five systems passed a majority of the tasks.

OpenAI reports that GPT-Rosalind passed 45.1 percent of text-only tasks but 28.1 percent of tasks containing artifacts or web links. The paper separately reports 44.5 percent for text-only tasks and 28.6 percent for tasks requiring attached artifacts, excluding the web-link grouping. Both comparisons point to the same weakness: extracting evidence from complex figures or large sequence files and carrying it into a complete decision was harder than responding from text alone.

Scientific charts and data files displayed for analysis (illustrative image)

Expert-written rubrics do not mean every answer was human-graded

The benchmark has meaningful independent review. OpenAI reports that 453 experts who did not write the tasks assessed their realism, scientific grounding and usefulness. The paper says accepted tasks also completed at least two rounds of expert review.

Grading is a separate issue. The paper says automated rubric scores were compared with independent expert scores on a stratified subset of model responses. It also discloses that automated or model-assisted grading, where used, applied the written rubrics. That check is stronger than unvalidated model judging, but it does not amount to human grading of every response in the reported ranking.

A whistle and score sheet represent review of benchmark grading (illustrative image)

OpenAI's dual role narrows the claim

The comparison included GPT-5.4, GPT-5.5, GPT-Rosalind, Gemini 3.1 Pro and Grok 4.3. It did not include an Anthropic model or a human-scientist baseline. The absence of either does not invalidate the reported scores, but it prevents the study from supporting claims about all leading AI systems or direct parity with researchers.

The authors explicitly disclose that OpenAI developed LifeSciBench and that the evaluated systems include OpenAI models, saying the results should be interpreted with that institutional context in mind. Outside replication would provide a stronger test of the ranking, especially if it used independently selected models, fully disclosed grading procedures and a relevant human comparison.

A scientist performs precise laboratory work beside analytical equipment (illustrative image)

The benchmark measures one exchange, not a research program

Each model received a task and its associated material once, then produced a final answer without clarification or iterative feedback. This standardized single-turn format makes model comparisons manageable, but research usually involves repeated experiments, revised hypotheses and responses to new evidence.

OpenAI says LifeSciBench is not a substitute for studying models in live research environments and does not measure downstream research impact. The benchmark supports a narrower conclusion: the tested models can produce useful partial work on some bounded tasks, while remaining unreliable on many tasks involving artifacts, exact outputs and operational constraints.

Laboratories need tests built around their own work

A biotechnology team considering these systems would learn more from testing its own figures, sequence files, data conventions and review procedures than from adopting the headline ranking as a procurement score. The relevant outcome is whether a model extracts the correct evidence, records uncertainty and produces output that survives expert review in that specific workflow.

The 36.1 percent result identifies substantial room for improvement, while the artifact gap identifies where failures concentrate. Neither number establishes productivity gains, scientific discovery or safe deployment. Those claims require independent evaluation in longer, iterative settings with auditable data and human oversight.

Frequently asked questions

Did GPT-Rosalind answer 36.1 percent of questions correctly?
Not in the ordinary meaning of accuracy. It met or exceeded 70 percent of the task-specific rubric points on 36.1 percent of tasks. Responses that fell below that threshold could still receive partial credit.

Why did tasks with artifacts produce lower pass rates?
The benchmark required models to extract and integrate information from materials such as figures, tables, PDFs, sequence files and web references. OpenAI's analysis says complex figures and large sequence files were recurring sources of difficulty.

Was LifeSciBench independently reviewed?
External experts reviewed the tasks and rubrics, and independent experts checked automated scoring on a stratified subset of responses. OpenAI still developed the benchmark, ran the comparison and published results that included its own models.

Can the ranking guide adoption in a biotechnology laboratory?
It can identify candidate capabilities and failure modes, but it does not replace testing on a laboratory's own data and workflow. The reported experiment was single-turn and did not measure performance across a live, iterative research program.