Xaira Therapeutics · Strategic Proposal
Open Frontier
Intelligence for
Drug Discovery
A Strategic Proposal for Xaira Leadership
Prepared by Bo Wang · Chief AI Scientist
July 2026
Landscape verified July 22, 2026
"We do not need to own general intelligence.
We need to own what general intelligence becomes inside Xaira."
Title slide. This is v2 of the strategic brief, expanded to 13 slides with appendix. Central thesis stated up front.
The recommendation
Xaira should own the
scientific intelligence layer
-
The open-weight frontier is accelerating. Frontier-level capability is spreading across multiple providers and becoming economically deployable.
-
Xaira does not need to build a general-purpose foundation model. Pretrained intelligence is not the bottleneck. Scientific grounding is.
-
Frontier models should be interchangeable infrastructure. Use the best available model for each task. Route intelligently. Replace as the landscape evolves.
-
Xaira owns the layer that transforms general intelligence into drug-discovery decisions. That layer — post-training, biological reasoning, X-Cell, experimental feedback — is the moat.
"We do not need to own general intelligence. We need to own what general intelligence becomes inside Xaira."
Set up the central recommendation clearly. This slide should take 2-3 minutes. The four bullets are the logical chain. The thesis is the conclusion.
Why now · Competitive landscape updated July 22, 2026
The open frontier is accelerating
| Model |
Developer |
Country |
Weights |
Context |
Key capability |
| DeepSeek-V4-Pro |
DeepSeek |
China |
Open weights · License review |
1M tokens |
MoE 1.6T/49B active · top open-weight reasoning |
| Qwen3-235B-A22B |
Alibaba / Qwen |
China |
Apache 2.0 |
131K (YaRN) |
MoE 235B/22B · thinking + non-thinking modes |
| GLM-5.2 |
Z.ai (Zhipu AI) |
China |
MIT |
1M tokens |
Long-horizon reasoning · agentic tool use |
| Kimi K2.6 |
Moonshot AI |
China |
Open weights · License review |
256K tokens |
Multimodal agentic · swarm orchestration |
| Llama 4 Maverick |
Meta |
USA |
Community License · Commercial OK |
1M tokens |
MoE 17B active/400B · native text + image |
| Inkling |
Thinking Machines Lab |
USA |
Acceptable Use Policy · Review |
1M tokens |
MoE 975B/41B · text + image + audio · calibrated confidence |
01 · Frontier capability spreading across multiple providers
02 · Long-context reasoning now widely available
03 · Agentic performance improving rapidly
04 · Multimodal capabilities entering open-weight systems
05 · Inference costs declining
06 · Organizations can now build proprietary intelligence on top
The question is no longer whether open-weight models will become capable enough. It is how Xaira turns rapidly improving general capabilities into proprietary scientific intelligence.
Models listed are NOT drug-discovery models — they are evidence the open frontier is moving fast.
Kimi K3 was announced but NOT released as of July 22, 2026; excluded from table.
Inkling is Thinking Machines Lab (correct name — NOT "Inking").
All license notes are advisory — legal review required before deployment.
Strategic case
Why open models are strategically attractive
-
Strategic independence. Reduce dependence on a small number of external providers and their roadmap decisions.
-
Data privacy. Self-hosted deployment keeps sensitive biological, chemical, clinical, and program data inside Xaira-controlled environments.
-
Domain specialization. Post-training, fine-tuning, RL, distillation, and retrieval can all be tuned to scientific workflows.
-
Deep integration. Tighter coupling with X-Cell, X-Scientist, internal data systems, lab infrastructure, and agentic workflows.
-
Model interchangeability. Modular architecture reduces switching costs as the landscape changes — though switching is never frictionless.
-
Cost control. More flexibility in balancing inference cost, latency, throughput, and model size — compare total cost of ownership, not just API price.
-
Continual learning. Open models provide a practical basis for controlled learning loops from Xaira experiments, expert feedback, and program outcomes.
Core economic argument
"Xaira can capture the progress of the entire open-model ecosystem while investing its proprietary resources in scientific intelligence that competitors cannot easily reproduce."
Seven distinct strategic reasons. Emphasize that model interchangeability is real but not frictionless — engineering work required. Cost argument is about TCO, not just API pricing.
Strategic balance
Open and closed models will both matter
Closed frontier models
API only
Strengths
Strongest capabilities on some tasks today
Mature managed infrastructure
Rapid access to new capabilities
Lower initial engineering burden
Limitations
Limited control over weights or training
Data-governance and privacy concerns
Changing pricing, access, and usage policies
Strategic dependency on provider roadmap
Open-weight models
Weights available
Strengths
Deployment control and private inference
Stronger customization and fine-tuning
Flexible integration with internal systems
Distillation and continual adaptation
Reduced provider dependency
Limitations
Infrastructure and talent requirements
Potentially weaker on some frontier tasks
Licensing complexity; security responsibilities
Recommendation: model-agnostic hybrid strategy
· Prefer open models where they meet scientific and operational requirements
· Retain closed models as teachers, evaluators, fallbacks, or specialized capability providers
· Route tasks based on performance, privacy, cost, modality, latency, and scientific risk
This is not an open-vs-closed argument. It is a hybrid architecture argument. Closed models remain valuable as teachers and evaluators even if they are not the primary inference backbone.
Geopolitical landscape
The global model ecosystem is diverging
🇨🇳 China
Open-weight acceleration
DeepSeek-V4-Pro
1.6T/49B active · 1M ctx · Text · DeepSeek (review license)
Qwen3-235B-A22B
235B/22B active · 131K ctx · Text · Apache 2.0
GLM-5.2
MoE · 1M ctx · Text · MIT license
Kimi K2.6
1T/32B active · 256K ctx · Text+Vision · (review license)
Themes: strong open-weight emphasis · rapid iteration · competitive inference economics · aggressive efficiency
Important considerations
License terms · commercial restrictions · cybersecurity review · supply-chain risk · model provenance · regulatory considerations
🇺🇸 United States
Closed frontier + emerging open
Closed frontier
GPT-5 / o3
OpenAI · API only · No weights · Strongest reasoning tasks
Claude Sonnet 4
Anthropic · API only · No weights · Strong safety + reasoning
Gemini 2.5 Pro
Google · API only · No weights · Multimodal frontier
Open-weight
Llama 4 Maverick
Meta · 17B active/400B · 1M ctx · Text+Image · Community License
Inkling
Thinking Machines Lab · 975B/41B · 1M ctx · Text+Image+Audio · AUP
Gap: US concentrating frontier capability in closed providers. Open ecosystem present but thinner than China's.
Two distinct trends. Chinese ecosystem strongly favoring open-weight releases with permissive licenses. US frontier capabilities concentrating in closed providers — with Meta and Thinking Machines Lab as exceptions. This divergence creates both opportunity and risk for Xaira.
Architecture strategy
Xaira should benefit from
both ecosystems
🇨🇳 Chinese open-weight
DeepSeek-V4-Pro · Qwen3 · GLM-5.2 · Kimi K2.6
+ future Chinese models
→
Xaira-Owned Layer
Scientific Intelligence
Post-training · Biological reasoning
X-Cell · Knowledge · Governance
→
Drug Discovery Decisions
Target selection · Experiment design
Program prioritization · Molecule design
↑
US open-weight (Llama 4, Inkling) + Closed frontier (GPT-5, Claude, Gemini) also feed in
"Xaira should not bet its strategy on which country or company wins the base-model race. It should build an architecture that benefits from progress across the entire frontier."
Data security requirement
Sensitive Xaira data must never be exposed to an external provider without appropriate legal, security, privacy, and technical review.
The key insight: Xaira should be agnostic about who wins the base-model race. The architecture should benefit from any improvement in the ecosystem. Data security is non-negotiable — emphasize this strongly to legal and security.
Strategic clarity
What Xaira should — and should not — own
—
A general-purpose pretrained model from scratch
—
The world's largest training cluster
—
Generic coding or language capabilities
—
A permanent commitment to one foundation model
●
Scientific post-training data and evaluation benchmarks
●
Biological and therapeutic reasoning
●
Program memory and internal knowledge
●
X-Cell and X-Scientist integration
●
Scientific tools and workflow orchestration
●
Expert feedback systems and experimental learning loops
●
Model routing and governance
●
The therapeutic decisions
"We do not need to own general intelligence. We need to own what general intelligence becomes inside Xaira."
This is the clarity slide. Make sure leadership leaves this slide understanding: we are NOT building a general model. We ARE building the scientific intelligence layer that uses any frontier model as interchangeable infrastructure.
Architecture
The Xaira scientific
intelligence architecture
Four layers. One feedback loop. The Xaira-owned layer is the core.
Frontier models are interchangeable. The scientific intelligence layer is the moat.
As better open-weight models appear, Xaira swaps the base layer — the proprietary stack above stays intact.
Xaira-owned layer (most prominent)
Interchangeable frontier layer
Xaira execution + outcomes
Interchangeable frontier model layer
Chinese open-weight · US open-weight · Closed frontier · Future models
◆ Xaira-owned scientific intelligence layer
Scientific reasoning · Domain post-training
Internal knowledge · Evidence attribution · Uncertainty calibration · Model routing · Governance
Xaira models and execution systems
X-Cell · Biological foundation models
X-Scientist · Scientific agents · Tools · Analysis pipelines · Lab systems
Therapeutic decisions and experiments
Target selection · Experiment design
Program prioritization · Molecule design · Translation strategy
The Xaira-owned layer (second from top) is highlighted with purple accent to make it the visual center. The feedback loop arrow on the right represents experiments and scientist decisions flowing back up into the intelligence layer — this is the learning loop that creates compounding value.
Xaira's moat
Why Xaira has the right to win
Most organizations can access the same frontier models. Very few can connect them to this.
🧬
Proprietary multimodal biological data
Internal omics, imaging, structural, experimental datasets
⚡
CRISPR perturbation data · X-Atlas
Millions of causal knockdown experiments
🔬
X-Cell
Causal virtual cell model trained on perturbation data
🤖
X-Scientist
Agentic scientific workflows at scale
🏥
Wet-lab infrastructure
Prospective experimental feedback loop
💊
Internal therapeutic programs
Real clinical context for grounding
🧑🔬
Multidisciplinary scientific expertise
Biology, chemistry, clinical — in the feedback loop
📡
Ability to generate new signal
Every assay, every experiment improves the system
Open models are the foundation. Xaira's scientific intelligence is the moat.
X-Cell and X-Scientist are highlighted with the Xaira accent. These are the core proprietary assets that make the architecture uniquely Xaira's. None of this is replicable by an AI lab releasing a general model.
Concrete first step
First pilot: target-validation
intelligence
One system. One workflow. Connected directly to Xaira's core drug-discovery process.
Why this pilot
✓ Connects directly to core drug-discovery process
✓ Demonstrates X-Cell integration
✓ Creates measurable feedback loop
✓ Clear decision utility for scientists
System capabilities
1
Ingest evidence
Internal and external evidence ingestion
2
Reason over X-Cell
Perturbation data and causal virtual cell
3
Generate hypotheses
Competing mechanistic hypotheses
4
Identify uncertainties
Critical unknowns that block decisions
5
Recommend experiments
Discriminating assays to resolve uncertainty
6
Coordinate and interpret
Agents + tools · Interpret results back into decision
7
Support the decision
Advance · modify · stop — with evidence
The target-validation pilot is not a demo — it is a real system connected to a real decision. The goal is the closed-loop learning system, not a proof-of-concept presentation.
Execution plan
Six-month roadmap
Foundation
-
·
Model evaluation suite design
-
·
Secure infrastructure setup
-
·
Team designation
-
·
License and legal review for candidate models
Build + Test
-
·
Target-validation intelligence pilot: build
-
·
X-Cell integration
-
·
Internal test with scientific team
-
·
Iteration based on feedback
Prospective Pilot
-
·
Deploy on active therapeutic program
-
·
Prospective evaluation against outcomes
-
·
First feedback loop measurement
-
·
Report back to leadership: scale, redirect, or stop
Start
Month 2
Month 4
Month 6 · Decision point
Six months is the minimum viable test. Month 6 is an explicit decision point — we are asking leadership to commit to six months, not to a permanent program. Scale, redirect, or stop.
The ask
We are asking for six months
to prove this works.
Approve the Open Frontier Intelligence initiative
Appoint an executive sponsor
Designate technical and scientific owner
Allocate a small senior cross-functional team
Establish secure model-evaluation infrastructure
Select the target-validation intelligence workflow as the flagship pilot
"We do not need to own general intelligence. We need to own what general intelligence becomes inside Xaira."
Decision at 6 months
Return to leadership with evidence.
Scale the program · redirect the focus · or stop.
This is a deliberate, bounded commitment.
This is the close. Be explicit: we are asking for six months and a small team. We commit to a clear evaluation at the end. Scale, redirect, or stop — leadership retains the decision.
Appendix A · Verified July 22, 2026
Detailed model specifications
| Model |
Developer |
Country |
Release |
Weights |
License |
Commercial |
Context |
Architecture |
Modalities |
| DeepSeek-V4-Pro |
DeepSeek |
China |
2026 |
Yes |
DeepSeek License |
Review |
1M tokens |
MoE 1.6T/49B active · CSA+HCA · 32T tokens trained |
Text |
| Qwen3-235B-A22B |
Alibaba / Qwen |
China |
2025 |
Yes |
Apache 2.0 |
Yes |
32K native / 131K YaRN |
MoE 235B/22B active · 128 experts · 94 layers |
Text |
| GLM-5.2 |
Z.ai (Zhipu AI) |
China |
Jun 13 2026 |
Yes |
MIT |
Yes |
1M tokens (128K out) |
MoE · IndexShare sparse attn · MTP |
Text |
| Kimi K2.6 |
Moonshot AI |
China |
2026 |
Yes |
Custom · Review |
Review |
256K tokens |
MoE 1T/32B active · 61 layers · MoonViT 400M |
Text · Vision |
| Llama 4 Maverick |
Meta |
USA |
Apr 5 2025 |
Yes |
Llama 4 Community |
Yes |
1M tokens |
MoE 17B active/400B total · 128 experts · early fusion |
Text · Image |
| Inkling |
Thinking Machines Lab |
USA |
Jul 2026 |
Yes |
Custom AUP |
Review |
1M tokens |
MoE 975B/41B active · 256 experts · 66 layers · hybrid local/global attn |
Text · Image · Audio |
Closed frontier models (API only — no weights)
| Model | Developer | Weights | Notes |
| GPT-5 / o3 |
OpenAI |
None |
Frontier reasoning · API only · data governance review required |
| Claude Sonnet 4 |
Anthropic |
None |
Strong safety + reasoning · API only · data governance review required |
| Gemini 2.5 Pro |
Google |
None |
Multimodal frontier · API only · data governance review required |
* Benchmark figures from model cards are developer self-reported and not independently verified. All license notes advisory only — legal review required before deployment.
Appendix. Full verified specs. Kimi K3 excluded — not released as of July 22, 2026. All license notes are advisory.
Appendix B · Legal review required
License details per model
Permissive
Apache 2.0 — Qwen3
Commercial use allowed. Attribution required. Patent rights granted. No trademark use of "Qwen." Standard Apache terms apply.
Permissive
MIT License — GLM-5.2
Commercial use allowed. Minimal restrictions. Retain copyright notice. Broadly permissive.
Commercial OK
Llama 4 Community License — Meta
Commercial use allowed. Restrictions apply for services with >700M monthly active users. Acceptable Use Policy must be followed. Available at github.com/meta-llama.
Review Required
DeepSeek Model License
Open weights available. License terms require review for commercial/internal deployment. Regulatory and supply-chain considerations apply for models from Chinese developers.
Review Required
Custom License — Kimi K2.6
Open weights available but custom license. Commercial and internal use terms require review. Verify on HuggingFace moonshotai/Kimi-K2.6 before deployment.
Review Required
Acceptable Use Policy — Inkling
Custom AUP from Thinking Machines Lab. NOT MIT or Apache. Weights downloadable but commercial deployment requires review of AUP terms. Safety evaluations conducted by developer.
All license information is advisory only. Consult legal before any deployment.
License review is mandatory before any deployment. Green = cleaner path. Warn = requires legal sign-off. Chinese models require additional cybersecurity and regulatory review beyond just the license terms.
Appendix C · If target-validation is not the right first pilot
Alternative pilot ideas
Scientific literature synthesis agent
Deploy a long-context open-weight model to synthesize internal and external scientific publications on a specific disease area or target class. Tests model integration and RAG pipeline without requiring X-Cell coupling.
Entry barrier: low · Decision utility: medium · X-Cell integration: no
Experiment design assistant
Build a system that takes a scientific question and proposes assay designs with justifications. Integrates with internal protocol databases and external literature. Validates against scientist judgment.
Entry barrier: medium · Decision utility: high · X-Cell integration: optional
Multi-model scientific reasoning benchmark
Before any deployment, run a structured evaluation of candidate models on proprietary Xaira scientific tasks. Outputs a ranked comparison informing all future model selection decisions.
Entry barrier: low · Decision utility: high (infrastructure) · X-Cell integration: optional
X-Atlas perturbation reasoning
Direct integration with X-Atlas CRISPR data. Ask open-weight models to reason over perturbation profiles, predict phenotypic consequences, and identify mechanistic hypotheses. Direct X-Cell coupling.
Entry barrier: medium-high · Decision utility: high · X-Cell integration: yes
Alternative pilots in case leadership wants to start with something lower risk before the full target-validation system. The multi-model benchmark is particularly good as a first step since it generates infrastructure value regardless of which pilot is chosen.
Appendix D · Verified July 22, 2026
Sources and verification notes
GLM-5.2
- HuggingFace: zai-org/GLM-5.2
- z.ai blog (release announcement)
- arXiv: 2602.15763
- License: MIT (verified on HuggingFace)
- Release: June 13, 2026 (verified)
Kimi K2.6
- HuggingFace: moonshotai/Kimi-K2.6
- Architecture: MoE 1T/32B active, 61 layers
- MoonViT: 400M params
- Kimi K3 status: announced, not released as of July 22, 2026
Inkling (Thinking Machines Lab)
- thinkingmachines.ai/inkling
- HuggingFace: thinkingmachines/Inkling
- Updated ~July 20, 2026
- License: Custom Acceptable Use Policy (NOT MIT)
- Note: correct name is "Inkling" not "Inking"
DeepSeek-V4-Pro
- HuggingFace: deepseek-ai/DeepSeek-V4-Pro
- arXiv: 2606.19348
- Training: 32T tokens, GRPO post-training
- License: DeepSeek Model License (review required)
Qwen3-235B-A22B
- HuggingFace: Qwen/Qwen3-235B-A22B
- qwenlm.github.io/blog/qwen3
- License: Apache 2.0 (verified from LICENSE file)
- Context: 32K native, 131K with YaRN
Llama 4 Maverick
- HuggingFace: meta-llama/Llama-4-Maverick-17B-128E-Instruct
- Release: April 5, 2025
- License: Llama 4 Community License Agreement
- Context: 1M tokens (Maverick), 10M (Scout)
All benchmark figures cited from model cards are developer self-reported. Web search unavailable during research — verified from HuggingFace model pages and official developer sites directly. Legal review of all licenses required before deployment.
Source appendix. All key claims traced to primary sources. No secondary sources used where avoidable. Kimi K3 excluded due to no public weights.