I built an AI agent team as biopharma executives, and sent them to the boardrooms where historical high-stakes strategic decisions took place, on a time machine.
In June 2012, both Claude Opus 4.6 and ChatGPT-5.3 AI agent teams, acting as BMS’s executive team, picked all-comers for the pivotal first-line lung cancer trial that would become CheckMate-026. It is the exact call that walked BMS into a failed Phase 3, a missed primary endpoint at HR 1.15, and the permanent loss of the first-line NSCLC market to Merck. Here is my previous article on this experiment, in case you haven’t read it before: Can AI Make Better Decisions Than Pharma Executives?
In April 2014, twenty-two months later, same five-agent AI team, same two models. I asked them to make Merck’s pivotal enrichment call for what would become KEYNOTE-024. Both models picked TPS ≥50% enrichment. That is the winning decision that made Keytruda the world’s best-selling drug, eventually clearing $25B in annual revenue.
Two decisions. Structurally identical: pick an enrichment strategy for a pivotal IO trial in 1L NSCLC. The models didn’t change. The framework didn’t change. The evidence base did.
That is the central finding of an experiment I extended from 1 scenario to 10 over the past quarter. Each one is a make-or-break historical biopharma decision, run independently through Claude Opus 4.6 and ChatGPT-5.3-Codex, with both models given identical time-locked mandates and no ability to actively inquire about anything that happened after the decision cutoff date. The full scorecard is below. But the BMS-to-Merck flip is the one I keep coming back to, because it tells you exactly what AI does well and exactly where it stops.
Then I ran the same framework forward on 10 open biopharma decisions whose answers won’t resolve until 2027 through 2031.
Recap: the original experiment
In the earlier piece Can AI Make Better Decisions Than Pharma Executives?, I asked a simple question: if today’s best AI agent team (CEO, R&D, Commercial, Regulatory, CI) ran the BMS leadership team’s CheckMate-026 decision, with the mandate to make decisions only with information available before June 2012, would the AI catch what the company missed?
It didn’t. Both Claude Opus 4.6 and Codex picked all-comers, just like the real BMS team. The result is uncomfortable but interesting: the AI failed in the same way humans failed, with the same data, under the same constraints. That is closer to a diagnostic than a defeat: it means the framework was working as intended (given that AI can technically “cheat” to get to the right decision if it relies on post-hoc correct information in pre-training data) and that the failure mode was the same one that defeats human executives.
Once I had a working harness, the natural next move was to scale it.
The scale-up: same framework, same models, 10 historical decisions
I picked 10 major historical biopharma decisions across different categories: trial design (biomarker, dose, indication), R&D go/no-go on safety signals, regulatory filing strategy, indication investment, and pricing. I ran each one independently through both Claude Opus 4.6 and GPT-5.3-Codex, then scored each output against the historical outcome.
Each scenario is a standalone setup with:
3-5 decision options the executive team was actually weighing at the time
An AI agent team of 5 that has to debate and converge. Each AI agent has a clean context window and clearly defined mandate (R&D head, Competitive Intel, Regulatory, Commercial). The CEO agent synthesizes all information, weighs pros and cons, and makes the final call.
Time-locked source instructions (PubMed, ClinicalTrials.gov, SEC EDGAR, Wayback Machine, press releases before time cutoff, and explicit prohibitions on using anything that postdates the cutoff)
Same framework as the original BMS piece. Same two models. Each output scored as Right (matched the historically vindicated path) or Wrong (matched the failed path). Where the AI chose a structurally similar but slightly different path that still produced the wrong outcome, I count it as Wrong. Half-credit hides the lesson.
A few honest disclaimers up front:
Models. I used Opus 4.6 and GPT-5.3 because those were the frontier when I started the initial retrospective run. Anthropic has since released Opus 4.7 and OpenAI released GPT-5.5, both meaningfully better at reasoning and tool use. I reran the methodology on the upgraded models for the forward-looking section below. The headline finding (the framework dominates the model) holds.
Sample. 10 scenarios is small. Individual scoring calls are debatable. Treat the numbers below as directional, not definitive. (I am currently running an upgraded pipeline with 40 more scenarios and many permutations on workflow setup and different contexts. If you are interested, let’s have a chat!)
The hindsight critique. The obvious objection: “Of course the AI gets these right: the outcome is in the training data”. The time-locked role-play is the explicit answer to that, and I’ll show below that the failures are inconsistent with pure outcome retrieval. If the AI were just looking up what happened, the three scenarios both models clearly missed (BMS 2012, Biogen aducanumab, Bluebird 2020) are the ones it would have nailed.
The scorecard: 10 historical decisions
1. BMS + CheckMate-026 enrichment (June 2012). Should BMS enrich for PD-L1 expression in the pivotal first-line NSCLC trial of nivolumab, or enroll all-comers?
Company actual decision: All-comers (no enrichment).
Right answer (in hindsight): Enrich for PD-L1 expression (≥5% or ≥50%).
Claude Opus 4.6 5-agent exec team: “All-comers. The 42-patient biomarker dataset is too thin to bet a $1B pivotal trial on.” (Wrong)
Codex GPT-5.3 5-agent exec team: “All-comers. The biomarker isn’t reliable enough yet.” (Wrong)
2. Merck + KEYNOTE-024 enrichment (April 2014). Should Merck enrich for PD-L1 TPS ≥50% in the pivotal first-line NSCLC trial of pembrolizumab, or use a lower threshold / all-comers?
Company actual decision: TPS ≥50% enrichment.
Right answer (in hindsight): TPS ≥50% enrichment (Keytruda became the world’s best-selling drug, >$25B/year).
Claude Opus 4.6 5-agent exec team: “TPS ≥50%. KEYNOTE-001 shows a clear response gradient. Lock in the contrarian bet.” (Right)
Codex GPT-5.3 5-agent exec team: “TPS ≥50%. Go for speed plus differentiation against BMS’s all-comers play.” (Right)
3. Daiichi Sankyo + HER2-low breast (December 2019). Should Daiichi go direct-to-Phase-3 for trastuzumab deruxtecan in HER2-low breast cancer (a patient population that didn’t formally exist) based on a 37% ORR in 54 Phase 1b patients and a clean bystander-effect mechanism?
Company actual decision: Direct Phase 3 (DESTINY-Breast04).
Right answer (in hindsight): Direct Phase 3 (PFS HR 0.50, OS HR 0.64; ~$4.4B projected revenue, new category created).
Claude Opus 4.6 5-agent exec team: “Direct Phase 3. The bystander mechanism is biologically clean enough.” (Right)
Codex GPT-5.3 5-agent exec team: “Direct Phase 3. First-mover advantage is the asymmetric upside.” (Right)
4. Pfizer + Torcetrapib ILLUMINATE (January 2005). Should Pfizer proceed to the 15,000-patient ILLUMINATE Phase 3 despite a 2-4 mmHg blood pressure signal in Phase 2, with torcetrapib positioned as Pfizer’s Lipitor successor ($12B/year line item)?
Company actual decision: Proceed directly to ILLUMINATE.
Right answer (in hindsight): Delay and investigate the BP signal mechanistically (ILLUMINATE was terminated for excess mortality, 82 vs 51 deaths; Pfizer lost ~$21B in market cap).
Claude Opus 4.6 5-agent exec team: “Proceed plus a parallel mechanistic BP investigation.” (Wrong: a parallel MoA study would not have caught the BP-to-mortality link in time)
Codex GPT-5.3 5-agent exec team: “Delay for a mechanistic BP study, run a backup CETP program in parallel.” (Right: would have likely avoided the disaster)
5. Biogen + aducanumab filing (June 2019). Should Biogen file aducanumab for FDA approval based on a post-hoc reanalysis after both Phase 3 trials (EMERGE, ENGAGE) were stopped for futility?
Company actual decision: File on the post-hoc reanalysis.
Right answer (in hindsight): Don’t file, or fully restart with a prespecified Phase 3 (Aduhelm got controversial AA approval over a 10-0 negative AdCom; CMS restricted coverage to trials only; ultimately withdrawn in Jan 2024).
Claude Opus 4.6 5-agent exec team: “File plus a binding confirmatory trial commitment.” (Wrong: the underlying program was structurally flawed; a confirmatory commitment would not have changed the AdCom vote or the CMS coverage block)
Codex GPT-5.3 5-agent exec team: “File plus a binding confirmatory trial commitment.” (Wrong: same reasoning)
6. Eisai + lecanemab Clarity AD (January 2021). Should Eisai proceed with the Clarity AD Phase 3 for lecanemab during the anti-amyloid nadir, immediately after the aducanumab AdCom debacle?
Company actual decision: Proceeded as designed.
Right answer (in hindsight): Proceed (Clarity AD met its primary endpoint at 27% CDR-SB slowing; first anti-amyloid to achieve traditional FDA approval).
Claude Opus 4.6 5-agent exec team: “Proceed. The biomarker-mechanism story is coherent, and the asymmetric payoff is large.” (Right)
Codex GPT-5.3 5-agent exec team: “Proceed with four explicit guardrails (ARIA monitoring, stopping rules, communication plan, post-marketing commitments).” (Right)
7. Intercept + OCA NASH filing (September 2019). Should Intercept file obeticholic acid for NASH broadly (F2-F3) or narrowly (F3-only), given that REGENERATE met the fibrosis endpoint but missed NASH resolution at 18 months, with 51% pruritus?
Company actual decision: Broad F2-F3 filing.
Right answer (in hindsight): F3-only filing, or not file at all (the broad filing earned a Complete Response Letter; Intercept lost first NASH approval to Madrigal’s resmetirom).
Claude Opus 4.6 5-agent exec team: “F3-only: narrower risk-benefit profile, better regulatory posture.” (Right)
Codex GPT-5.3 5-agent exec team: “F3-only: match label to the strongest subgroup signal.” (Right)
8. Novo Nordisk + semaglutide obesity STEP (January 2016). Should Novo invest $1-1.5B in a dedicated obesity Phase 3 program (STEP) for semaglutide 2.4mg, despite the obesity drug market being widely regarded as a “commercial wasteland”?
Company actual decision: Full STEP program.
Right answer (in hindsight): Full STEP program (Wegovy showed ~15% weight loss; obesity market exploded past $50B; Novo became one of the world’s most valuable companies).
Claude Opus 4.6 5-agent exec team: “Full STEP: 12-15% weight loss is a qualitative step-change in obesity efficacy.” (Right)
Codex GPT-5.3 5-agent exec team: “Full STEP: the competitive window is open and the payoff is asymmetric.” (Right)
9. Bluebird + Zynteglo EU pricing (January 2020). Should Bluebird set the European launch price and reimbursement model for Zynteglo at roughly €1.5-1.8M with 5-year outcomes-based installments?
Company actual decision: €1.575M with 5-year outcomes-based installments.
Right answer (in hindsight): Failure. Any version of this pricing model was likely doomed because of EU HTA structural barriers (Bluebird exited the European market in August 2021).
Claude Opus 4.6 5-agent exec team: “$1.8M with installments. QALY-based value defensible, orphan benchmarks support.” (Wrong)
Codex GPT-5.3 5-agent exec team: “$1.575M with installments. Outcomes-based contracting precedent supports.” (Wrong)
10. Sarepta + Elevidys filing (January 2022). Should Sarepta file Elevidys gene therapy for DMD despite missing the functional endpoint, leaning on biomarker (micro-dystrophin) data?
Company actual decision: Pre-BLA meeting → narrow filing (ages 4-5).
Right answer (in hindsight): Pre-BLA meeting → narrow filing (AdCom voted 8–6 in favor of accelerated approval, May 2023; FDA approved June 2023 at $3.2M for ages 4–5; confirmatory data has continued to disappoint, a “mixed” outcome).
Claude Opus 4.6 5-agent exec team: “Pre-BLA meeting then narrow filing: gain the time-window without overreaching.” (Right)
Codex GPT-5.3 5-agent exec team: “Pre-BLA meeting then narrow filing for the youngest subgroup.” (Right)
Aggregate:
Claude Opus 4.6 5-agent team: 6 Right, 4 Wrong (60%)
Codex GPT-5.3 5-agent team: 7 Right, 3 Wrong (70%)
Head-to-head: Both models agreed on the action in 9 of 10 scenarios; they only diverged on Torcetrapib, where Codex’s call would have likely saved Pfizer roughly $21B.
The numbers are useful but the patterns are more interesting.
Pattern 1. The consensus floor is high
The most consistent finding across the dataset: when the historically vindicated decision was also the consensus view of the period’s experts, both AIs almost always got it. Merck’s TPS ≥50% bet was supported by a clear PD-L1 response gradient in KEYNOTE-001. Daiichi’s HER2-low Phase 3 leap had a 37% ORR signal in 54 patients plus a clean bystander-effect mechanism. Eisai’s lecanemab Clarity AD push had a coherent biomarker-and-mechanism story even in the anti-amyloid nadir. Novo’s obesity bet had STEP-program-level weight loss already visible in the Phase 2.
In all of those, both models converged on the historically correct call, often with moderate-to-high confidence and unanimous internal “votes” from the five “expert agents.” That is the floor of AI capability on these decisions, which might be meaningfully higher than some executive teams hit in practice. Real executives carry sunk cost, internal politics, quarter-end pressure, and motivated reasoning about their own programs. The five-agent role-play has none of those incentives. It just synthesizes the available evidence.
This is the most underrated finding in the dataset. The discourse on “AI in the C-suite” obsesses over the spectacular cases. Will it find the contrarian gem? Will it spot the fraud? But the boring version sometimes matters more.
Pattern 2. The contrarian ceiling is low
Now the other side. The two failures that fell short of an actual contrarian call (Scenario 1 BMS PD-L1, Scenario 9 Bluebird EU pricing) share one feature: the right answer required betting against the prevailing view with thin or absent direct evidence.
The BMS case is the cleanest illustration. In June 2012, the entire field’s biomarker dataset for PD-L1 in NSCLC was about 42 patients. The internal modal view was “we need to study all-comers, the biomarker isn’t reliable enough.” 24 months later, with hundreds more patients, the modal view at Merck flipped. Both AIs followed the modal view at each time point. Neither model, in 2012, made the leap that Merck would make in 2014, because the 2014 evidence didn’t exist yet.
This is the part that should be uncomfortable for anyone deploying AI as a strategic decision-maker. The Keytruda vs. Opdivo comparison is the proof. Same model. Same framework. Two years apart. Both calls perfectly aligned with the consensus of the time. The AI is, fundamentally, a consensus reasoner. That is a good thing when consensus is correct, and a terrible thing when the contrarian play is the operator’s edge.
Pattern 3. The two models barely disagree
The one disagreement is Torcetrapib: Claude said 'proceed plus parallel investigation,' Codex said 'delay and investigate.' The latter is the call that would have saved Pfizer ~$21B.
Beyond that single case: different lab, different RLHF, different tokenizer; same answer 90% of the time. On the cleanest correct calls (Merck enrichment, Enhertu HER2-low, Novo STEP, Eisai lecanemab), the agreement was unanimous. On the cleanest failures (BMS PD-L1, Bluebird EU, Biogen aducanumab), the agreement was also unanimous. Both walked into the same trap with the same reasoning.
This matters for two reasons. First, it suggests the agent framework is doing more of the work than the underlying model. The five-agent role-play, the time-locked source instructions, the structured option set, and the requirement to converge on a single decision. Those are the load-bearing elements. Swap the model out and the answer doesn’t move much. That has a corollary: if you want a better AI decision-maker, working on the framework is maybe higher leverage than waiting for the next model release.
Second, the consensus across two frontier models with different training methodologies is a useful signal in itself. When Claude and Codex independently arrive at the same answer with the same reasoning, that is at least directionally meaningful evidence that the framework’s output reflects something real about the period’s available evidence, not a model-specific bias. When they disagree (as they did on Torcetrapib), that is exactly the kind of judgment call worth escalating to a human.
What this means for the decision makers
Use AI as a check, not a decider. The consensus floor is high enough that the framework is a useful sanity check on any non-trivial strategic decision. If the AI converges on the same answer your team is leaning toward, you might have eliminated one class of error. If it converges on a different answer, that is the meeting you should actually have.
Lean on AI hardest where the decision turns on data interrogation. Codex’s call on Torcetrapib and both AIs’ call on Intercept OCA-NASH suggest the framework is more risk-aware than the typical leadership team under commercial pressure. The five-agent structure forces the regulatory voice to be present at full weight, every time.
Don’t expect contrarian conviction. The framework tracks the period’s consensus. If your operator edge depends on a contrarian bet, you have to bring that conviction yourself. The AI will give you a competent version of the consensus view, which is sometimes what you want and sometimes the opposite of what you need.
The framework matters more than the model. Of the variance in AI output across this dataset, very little of it is model-driven and most of it is framework-driven. The investment that compounds is in better role-play structures, better time-locking, better source instructions, better option-set framing. The investment that depreciates is in a particular model snapshot.
Part 2: I ran the framework forward. Ten open decisions, new models, harder test
Everything above is retrospective. The “the model knows the answer” critique only fully evaporates when you point the framework at decisions where the answer doesn’t exist yet. So I ran it forward: ten open biopharma decisions, upgraded models (Claude Opus 4.7 and Codex GPT-5.5), and an upgraded role-play (added CFO, CMO, and a Contrarian Board Member explicitly mandated to find asymmetric positions). Each CEO now commits to a single calibrated probability of success.
1. Novo + Metsera (September 22, 2025). Should Novo Nordisk submit a competing topping bid for Metsera Pharmaceuticals after Pfizer’s accepted $4.9B offer for the monthly-dosed GLP-1 (MET-097i)?
Company actual decision: Novo submitted competing topping bids through October–November 2025, reaching ~$10B. Metsera deemed Novo’s October 30 proposal a “Superior Company Proposal,” but the FTC raised antitrust concerns about Novo’s structure. On November 6 Novo declined to bid further; on November 7 Metsera accepted Pfizer’s amended ~$10B offer ($65.60 cash + $20.65 CVR). Pfizer prevailed.
Claude Opus 4.7 7-agent exec team: “Submit a superior proposal at ~$11.5B effective ($47.50 cash + $33.50 CVR), with a $750M reverse termination fee and pre-negotiated FTC divestiture undertaking.” (Confidence: 43%)
Codex GPT-5.5 7-agent exec team: “Submit a CVR-heavy, price-disciplined counter rather than a full $16.5B all-in topping bid. Walk away above $13B.” (Confidence: 32%)
2. BMS + Karuna (December 22, 2023). Should BMS acquire Karuna Therapeutics for $14B for KarXT (xanomeline-trospium), the first non-dopamine antipsychotic, as a hedge against the looming patent cliff?
Company actual decision: BMS proceeded with the full $14B acquisition; deal closed March 2024. KarXT approved by FDA September 2024 as Cobenfy.
Claude Opus 4.7 7-agent exec team: “Yes, proceed with the full $14B acquisition at $330/share.” (Confidence: 43%)
Codex GPT-5.5 7-agent exec team: “Yes, proceed with the full $14B acquisition at $330/share.” (Confidence: 57%)
3. Vertex + Alpine (April 10, 2024). Should Vertex acquire Alpine Immune Sciences for $4.9B to enter immunology via povetacicept (APRIL/BAFF inhibitor) for IgA nephropathy?
Company actual decision: Vertex proceeded with the full $4.9B acquisition; deal closed mid-2024.
Claude Opus 4.7 7-agent exec team: “Yes, proceed with the full $4.9B acquisition at $65/share.” (Confidence: 32%)
Codex GPT-5.5 7-agent exec team: “Yes, proceed with the full $4.9B acquisition at $65/share.” (Confidence: 32%)
4. BMS + RayzeBio (December 26, 2023). Should BMS acquire RayzeBio for $4.1B (fifteen days after closing the $14B Karuna deal) for an actinium-225 radioligand therapy platform?
Company actual decision: BMS proceeded with the full $4.1B acquisition; deal closed Q1 2024.
Claude Opus 4.7 7-agent exec team: “Yes, proceed with the full $4.1B acquisition at $62.50/share, with platform-entry tripwires (Ac-225 supply redundancy, ACTION-1 readout, follow-on program proof).” (Confidence: 33%)
Codex GPT-5.5 7-agent exec team: “Yes, proceed with the full $4.1B acquisition. Treat as scarce platform entry not asset certainty.” (Confidence: 30%)
5. Lilly + Morphic (July 8, 2024). Should Lilly acquire Morphic Therapeutic for $3.2B for MORF-057, an oral α4β7 integrin inhibitor for IBD?
Company actual decision: Lilly proceeded with the full $3.2B acquisition at $57/share; deal closed August 2024.
Claude Opus 4.7 7-agent exec team: “Yes, proceed with the full $3.2B acquisition.” (Confidence: 42%)
Codex GPT-5.5 7-agent exec team: “Yes, proceed with the full $3.2B acquisition at $57/share.” (Confidence: 32%)
6. Amgen MariTide Phase 3 (November 26, 2024). Should Amgen pursue an aggressive head-to-head Phase 3 of MariTide (monthly GLP-1) against tirzepatide, partner ex-US, or run a smaller program?
Company actual decision: Amgen advanced an aggressive monotherapy Phase 3 program (MARITIME: 72-week, three target doses 21/35/70mg) without an ex-US partner. Also initiated Phase 3 in ASCVD, heart failure, and obstructive sleep apnea in 2025.
Claude Opus 4.7 7-agent exec team: “Hybrid. Amgen retains US rights with full economics; negotiate an ex-US partnership at $2-3B upfront.” (Confidence: 33%)
Codex GPT-5.5 7-agent exec team: “Staged differentiated Phase 3. Emphasize off-drug durability claim under a phased dose program.” (Confidence: 28%)
7. AZ + Daiichi Sankyo: Dato-DXd filing scope (January 15, 2024). Should AstraZeneca/Daiichi file Dato-DXd in all-comers NSCLC despite a squamous subgroup harm signal (HR 1.32), or restrict the label to non-squamous?
Company actual decision: Initially filed non-squamous NSCLC BLA in February 2024. After FDA feedback, voluntarily withdrew the NSCLC BLA in Q4 2024 and refiled for the narrower EGFR-mutated NSCLC population, approved as Datroway in January 2025. Separately got HR+ HER2-low/negative breast cancer approval (Datroway, January 2025).
Claude Opus 4.7 7-agent exec team: “Restrict. File non-squamous NSCLC plus HR+ HER2-low/0 breast cancer; protect the narrative.” (Confidence: 28%)
Codex GPT-5.5 7-agent exec team: “Restrict. File non-squamous NSCLC plus HR+ HER2-low/0 breast cancer; explicit ILD risk management.” (Confidence: 34%)
8. AbbVie emraclidine (November 11, 2024). Should AbbVie continue Phase 3 of emraclidine after both Phase 2 trials missed primary endpoints, eleven months after acquiring Cerevel for $8.7B?
Company actual decision: AbbVie took a ~$3.5B impairment charge (January 2025 SEC filing) and did not advance Phase 3 in monotherapy schizophrenia. Pivoted the program to adjunctive schizophrenia therapy plus neurodegenerative psychosis (Alzheimer’s-associated psychosis), directionally close to what both AIs recommended.
Claude Opus 4.7 7-agent exec team: “Discontinue AbbVie-led monotherapy schizophrenia Phase 3; pivot to CIAS and AD-psychosis Phase 2 with redesigned methodology.” (Confidence: 20%)
Codex GPT-5.5 7-agent exec team: “Pause acute-schizophrenia Phase 3; pivot to a tightly gated salvage program in CIAS/negative symptoms or AD neuropsychiatric symptoms, only after failure forensics and FDA input.” (Confidence: 27%)
9. Vertex + Casgevy launch (December 15, 2023). Should Vertex launch Casgevy (the first FDA-approved CRISPR gene therapy for sickle cell disease) aggressively (50 ATCs, 1000+ patients in Year 1), or run a measured launch?
Company actual decision: Measured launch. Vertex prioritized ATC qualification quality over speed. ~50 ATCs authorized by end of 2024, with patient throughput limited by manufacturing capacity (low double-digits dosed in Year 1 globally).
Claude Opus 4.7 7-agent exec team: “Modified measured launch. $2.2M list pricing, 25 ATCs by end-2024, accelerated Middle East rollout, Q2 2024 opt-in to CMS Cell and Gene Therapy Access Model.” (Confidence: 30%)
Codex GPT-5.5 7-agent exec team: “Measured launch. 20-25 high-quality ATCs through 2024; preserve $2.2M WAC, use net-price flexibility via outcomes-based Medicaid contracts.” (Confidence: 20%)
10. BridgeBio Attruby pricing (November 22, 2024). Should BridgeBio price Attruby (acoramidis) for transthyretin amyloid cardiomyopathy at parity with Pfizer’s tafamidis ($268K/year), undercut, or price at a premium?
Company actual decision: Parity-minus list price of $244,500/year (~9% discount to Pfizer tafamidis at $268K WAC); launched Q1 2025 with strong commercial uptake; Q1 2025 net revenue beat sell-side expectations.
Claude Opus 4.7 7-agent exec team: “Parity-minus at $244,500/year plus 150-250 cardiology specialty reps, outcomes-based contracts, IRA negotiation positioning.” (Confidence: 45%)
Codex GPT-5.5 7-agent exec team: “Parity-minus/access-first launch at $245K-$250K; high-touch patient services and disciplined comparative messaging.” (Confidence: 50%)
The headline: Both models picked the same action on 9 of 10 scenarios; aggregate means were nearly identical (Claude 0.351, Codex 0.354). Per-scenario probabilities differ by 7 points on average; only Vertex-Alpine matched to two decimals. Three findings worth pulling out:
Option-level convergence holds; probability calibration is model-dependent. The retrospective set gave 9 of 10 ties on the action. The forward set, with harder conditions (newer models, more roles, no known answer), gives 9 of 10 too. The framework now forces a calibrated probability commitment, on which the two models clearly differ. The framework dominates on direction. It does not dominate on magnitude.
Where this leaves the experiment
From n=1 to n=20 cases, AI “consensus convergence” is more prominent. In over 90% of cases, Claude and GPT/codex agree with each other, surviving the model upgrade (4.6→4.7, 5.3→5.5), the role expansion (5 seats → 7), and the regime change from retrospective to forward-looking. Two models with similar overall risk appetite, but slightly different per-scenario calibration..
The thing I still keep coming back to is the BMS-to-Merck flip. Same model, same framework, twenty-two months apart, opposite answer. The evidence base bent, and the AI followed it. The operators who win these calls bend the evidence base before it bends them.
If we want to use AI to facilitate decision-making, instead of chasing the ever-changing SOTA models, we should think harder about agent harness design, evidence gathering, decision-making framework setup, and risk-reward appetite definition.
And more importantly, AI can give you the best possible analysis. It can’t give you the courage to go against it.














Very cool! How much context have you used as input for the LLM for each decision? Would be interesting to test it under different context/assumptions to see how the decisions would change!
Fascinating experiment. The future of AI won’t just be smarter models, it’ll be how well multiple agents coordinate, validate outputs, and handle real-world complexity together.