GPT-5 Pro generated a mechanism for unexplained T-cell data and predicted the direction of an unpublished CAR-T experiment, according to a case study released in November 2025. The result shows a model helping an expert turn laboratory observations into experiments, but the public record does not support treating the proposed mechanism as a confirmed discovery.
The case appears in an arXiv paper co-authored by Derya Unutmaz of The Jackson Laboratory and researchers from OpenAI and several universities and laboratories. It is one of a curated group of examples across biology, mathematics, physics and computer science, rather than a systematic evaluation with a representative sample of successes and failures.
The model proposed a mechanism beyond glucose restriction
Unutmaz's laboratory had treated human CD4-positive T cells with varying amounts of 2-deoxy-D-glucose (2-DG), a glucose analog that interferes with glucose metabolism. After two days, researchers removed the treatment and expanded the cells with interleukin-2 (IL-2). The cells later showed a persistent shift toward a proinflammatory Th17-like state, but the laboratory did not have a clear mechanism.
The paper says GPT-5 Pro analyzed an unpublished flow-cytometry figure in 17 minutes and proposed that 2-DG had impaired N-linked glycosylation, weakening IL-2 signaling and releasing a constraint on Th17 differentiation. It also predicted that memory T cells, rather than naive T cells, accounted for the effect. These were mechanistic hypotheses derived from the figure and established biological knowledge, not measurements made by the model.
A hidden result provided a stronger test
The researchers then supplied another unpublished figure involving CD8-positive naive and memory T cells. The model interpreted changes in the checkpoint proteins PD-1 and LAG-3 after transient 2-DG exposure, attributing them to impaired glycosylation and reduced T-cell receptor signaling.
When asked to simulate an anti-CD19 CAR-T experiment, GPT-5 Pro predicted that prior 2-DG treatment would increase the engineered cells' killing of target cancer cells, and the authors report that this direction nearly matched an experiment the laboratory had already run but not published. A result withheld from the prompt is a better test than asking for a plausible explanation alone because it creates a comparison against information the model was not explicitly given.
Unpublished does not guarantee unseen
The novelty claim has a documented limit. The model suggested a mannose-rescue experiment intended to restore glycosylation without restoring glycolysis, and its prediction matched work the laboratory had already conducted. Yet the authors disclose that a similar finding, including a mannose-rescue experiment, had previously appeared in a bioRxiv preprint.
The authors therefore say GPT-5 Pro may have encountered that finding and connected it to the new figure. They identify other proposed tests involving glycosylation inhibitors and the IL-2 pathway as unperformed and unpublished in this setting. Those experiments could distinguish a useful hypothesis from a mechanism that merely fits the available observations.
The evidence supports a workflow, not an autonomous scientist
The case depended on an experienced immunologist selecting the question, providing interpretable data and judging whether the model's biological reasoning deserved a test. The laboratory also held results that could be kept outside the prompt and used for comparison. Without those steps, a fluent mechanism would remain difficult to separate from a plausible error.
OpenAI describes the paper's examples as curated illustrations rather than a systematic sample and says GPT-5 can hallucinate citations, mechanisms and proofs. The company also says the model does not independently run research projects. That disclosure prevents this case from establishing an overall success rate or showing how the method performs with noisy data, weaker prompts or researchers outside the original collaboration.
Validation remains the scarce step
The clearest contribution was speed at an early stage of inquiry. GPT-5 Pro connected an observed phenotype with a possible pathway and produced experiments that could falsify or refine that explanation. It did not collect the samples, establish the provenance of every idea or complete the experiments needed to support the mechanism.
For research organizations, the case points to a three-part process: controlled access to well-described data, domain experts who can reject biologically weak output, and experiments reserved for validation. A more capable model may generate candidates faster, but credibility still comes from how those candidates are tested and reported.
Frequently asked questions
Did GPT-5 Pro solve the T-cell mechanism?
No. It proposed an explanation involving N-linked glycosylation and IL-2 signaling and suggested experiments. The paper says some of those experiments still had not been performed or published, so the mechanism remained a hypothesis.
What did the unpublished CAR-T result show?
The authors say the model predicted that a brief 2-DG exposure would enhance the killing activity of anti-CD19 CAR-T cells and that the prediction nearly matched their unpublished experiment. The underlying result was not available for independent inspection in the cited paper.
Could the model have recalled the answer from training data?
That possibility cannot be excluded for the mannose-rescue suggestion because a related experiment had appeared in a bioRxiv preprint. The specific figures used in the interaction and the CAR-T experiment were described as unpublished, but their status alone cannot prove that every relevant concept was absent from training data.