Xaira Therapeutics · Confidential

Xaira Therapeutics · Strategic Proposal

Open Frontier
Intelligence for
Drug Discovery

A Strategic Proposal for Xaira Leadership

Prepared by  Bo Wang · Chief AI Scientist
July 2026
Landscape verified July 22, 2026

"We do not need to own general intelligence.
We need to own what general intelligence becomes inside Xaira."

Title slide. This is v2 of the strategic brief, expanded to 13 slides with appendix. Central thesis stated up front.

The recommendation

Xaira should own the
scientific intelligence layer

  • The open-weight frontier is accelerating. Frontier-level capability is spreading across multiple providers and becoming economically deployable.
  • Xaira does not need to build a general-purpose foundation model. Pretrained intelligence is not the bottleneck. Scientific grounding is.
  • Frontier models should be interchangeable infrastructure. Use the best available model for each task. Route intelligently. Replace as the landscape evolves.
  • Xaira owns the layer that transforms general intelligence into drug-discovery decisions. That layer — post-training, biological reasoning, X-Cell, experimental feedback — is the moat.

"We do not need to own general intelligence. We need to own what general intelligence becomes inside Xaira."

Set up the central recommendation clearly. This slide should take 2-3 minutes. The four bullets are the logical chain. The thesis is the conclusion.

Why now · Competitive landscape updated July 22, 2026

The open frontier is accelerating

Model Developer Country Weights Context Key capability
DeepSeek-V4-Pro DeepSeek China Open weights · License review 1M tokens MoE 1.6T/49B active · top open-weight reasoning
Qwen3-235B-A22B Alibaba / Qwen China Apache 2.0 131K (YaRN) MoE 235B/22B · thinking + non-thinking modes
GLM-5.2 Z.ai (Zhipu AI) China MIT 1M tokens Long-horizon reasoning · agentic tool use
Kimi K2.6 Moonshot AI China Open weights · License review 256K tokens Multimodal agentic · swarm orchestration
Llama 4 Maverick Meta USA Community License · Commercial OK 1M tokens MoE 17B active/400B · native text + image
Inkling Thinking Machines Lab USA Acceptable Use Policy · Review 1M tokens MoE 975B/41B · text + image + audio · calibrated confidence
01 · Frontier capability spreading across multiple providers
02 · Long-context reasoning now widely available
03 · Agentic performance improving rapidly
04 · Multimodal capabilities entering open-weight systems
05 · Inference costs declining
06 · Organizations can now build proprietary intelligence on top

The question is no longer whether open-weight models will become capable enough. It is how Xaira turns rapidly improving general capabilities into proprietary scientific intelligence.

Models listed are NOT drug-discovery models — they are evidence the open frontier is moving fast. Kimi K3 was announced but NOT released as of July 22, 2026; excluded from table. Inkling is Thinking Machines Lab (correct name — NOT "Inking"). All license notes are advisory — legal review required before deployment.

Strategic case

Why open models are strategically attractive

  • Strategic independence. Reduce dependence on a small number of external providers and their roadmap decisions.
  • Data privacy. Self-hosted deployment keeps sensitive biological, chemical, clinical, and program data inside Xaira-controlled environments.
  • Domain specialization. Post-training, fine-tuning, RL, distillation, and retrieval can all be tuned to scientific workflows.
  • Deep integration. Tighter coupling with X-Cell, X-Scientist, internal data systems, lab infrastructure, and agentic workflows.
  • Model interchangeability. Modular architecture reduces switching costs as the landscape changes — though switching is never frictionless.
  • Cost control. More flexibility in balancing inference cost, latency, throughput, and model size — compare total cost of ownership, not just API price.
  • Continual learning. Open models provide a practical basis for controlled learning loops from Xaira experiments, expert feedback, and program outcomes.

Core economic argument

"Xaira can capture the progress of the entire open-model ecosystem while investing its proprietary resources in scientific intelligence that competitors cannot easily reproduce."

Seven distinct strategic reasons. Emphasize that model interchangeability is real but not frictionless — engineering work required. Cost argument is about TCO, not just API pricing.

Strategic balance

Open and closed models will both matter

Closed frontier models API only
Strongest capabilities on some tasks today
Mature managed infrastructure
Rapid access to new capabilities
Lower initial engineering burden
External API dependency
Limited control over weights or training
Data-governance and privacy concerns
Changing pricing, access, and usage policies
Strategic dependency on provider roadmap
Open-weight models Weights available
Deployment control and private inference
Stronger customization and fine-tuning
Flexible integration with internal systems
Distillation and continual adaptation
Reduced provider dependency
Infrastructure and talent requirements
Potentially weaker on some frontier tasks
Licensing complexity; security responsibilities
Rapid model obsolescence
Recommendation: model-agnostic hybrid strategy
· Prefer open models where they meet scientific and operational requirements
· Retain closed models as teachers, evaluators, fallbacks, or specialized capability providers
· Route tasks based on performance, privacy, cost, modality, latency, and scientific risk
This is not an open-vs-closed argument. It is a hybrid architecture argument. Closed models remain valuable as teachers and evaluators even if they are not the primary inference backbone.

Geopolitical landscape

The global model ecosystem is diverging

🇨🇳 China Open-weight acceleration
DeepSeek-V4-Pro
1.6T/49B active · 1M ctx · Text · DeepSeek (review license)
Qwen3-235B-A22B
235B/22B active · 131K ctx · Text · Apache 2.0
GLM-5.2
MoE · 1M ctx · Text · MIT license
Kimi K2.6
1T/32B active · 256K ctx · Text+Vision · (review license)
Themes: strong open-weight emphasis · rapid iteration · competitive inference economics · aggressive efficiency

Important considerations

License terms · commercial restrictions · cybersecurity review · supply-chain risk · model provenance · regulatory considerations

🇺🇸 United States Closed frontier + emerging open
Closed frontier
GPT-5 / o3
OpenAI · API only · No weights · Strongest reasoning tasks
Claude Sonnet 4
Anthropic · API only · No weights · Strong safety + reasoning
Gemini 2.5 Pro
Google · API only · No weights · Multimodal frontier
Open-weight
Llama 4 Maverick
Meta · 17B active/400B · 1M ctx · Text+Image · Community License
Inkling
Thinking Machines Lab · 975B/41B · 1M ctx · Text+Image+Audio · AUP
Gap: US concentrating frontier capability in closed providers. Open ecosystem present but thinner than China's.
Two distinct trends. Chinese ecosystem strongly favoring open-weight releases with permissive licenses. US frontier capabilities concentrating in closed providers — with Meta and Thinking Machines Lab as exceptions. This divergence creates both opportunity and risk for Xaira.

Architecture strategy

Xaira should benefit from
both ecosystems

🇨🇳 Chinese open-weight

DeepSeek-V4-Pro · Qwen3 · GLM-5.2 · Kimi K2.6
+ future Chinese models

Xaira-Owned Layer

Scientific Intelligence

Post-training · Biological reasoning
X-Cell · Knowledge · Governance

Drug Discovery Decisions

Target selection · Experiment design
Program prioritization · Molecule design

US open-weight (Llama 4, Inkling) + Closed frontier (GPT-5, Claude, Gemini) also feed in

"Xaira should not bet its strategy on which country or company wins the base-model race. It should build an architecture that benefits from progress across the entire frontier."

Data security requirement

Sensitive Xaira data must never be exposed to an external provider without appropriate legal, security, privacy, and technical review.

The key insight: Xaira should be agnostic about who wins the base-model race. The architecture should benefit from any improvement in the ecosystem. Data security is non-negotiable — emphasize this strongly to legal and security.

Strategic clarity

What Xaira should — and should not — own

✗ Xaira does NOT need to own
A general-purpose pretrained model from scratch
The world's largest training cluster
Generic coding or language capabilities
A permanent commitment to one foundation model
✓ Xaira SHOULD own
Scientific post-training data and evaluation benchmarks
Biological and therapeutic reasoning
Program memory and internal knowledge
X-Cell and X-Scientist integration
Scientific tools and workflow orchestration
Expert feedback systems and experimental learning loops
Model routing and governance
The therapeutic decisions

"We do not need to own general intelligence. We need to own what general intelligence becomes inside Xaira."

This is the clarity slide. Make sure leadership leaves this slide understanding: we are NOT building a general model. We ARE building the scientific intelligence layer that uses any frontier model as interchangeable infrastructure.

Architecture

The Xaira scientific
intelligence architecture

Four layers. One feedback loop. The Xaira-owned layer is the core.

Frontier models are interchangeable. The scientific intelligence layer is the moat. As better open-weight models appear, Xaira swaps the base layer — the proprietary stack above stays intact.

Xaira-owned layer (most prominent)
Interchangeable frontier layer
Xaira execution + outcomes
FEEDBACK LOOP
Interchangeable frontier model layer
Chinese open-weight · US open-weight · Closed frontier · Future models
◆ Xaira-owned scientific intelligence layer
Scientific reasoning · Domain post-training
Internal knowledge · Evidence attribution · Uncertainty calibration · Model routing · Governance
Xaira models and execution systems
X-Cell · Biological foundation models
X-Scientist · Scientific agents · Tools · Analysis pipelines · Lab systems
Therapeutic decisions and experiments
Target selection · Experiment design
Program prioritization · Molecule design · Translation strategy
The Xaira-owned layer (second from top) is highlighted with purple accent to make it the visual center. The feedback loop arrow on the right represents experiments and scientist decisions flowing back up into the intelligence layer — this is the learning loop that creates compounding value.

Xaira's moat

Why Xaira has the right to win

Most organizations can access the same frontier models. Very few can connect them to this.

🧬
Proprietary multimodal biological data
Internal omics, imaging, structural, experimental datasets
CRISPR perturbation data · X-Atlas
Millions of causal knockdown experiments
🔬
X-Cell
Causal virtual cell model trained on perturbation data
🤖
X-Scientist
Agentic scientific workflows at scale
🏥
Wet-lab infrastructure
Prospective experimental feedback loop
💊
Internal therapeutic programs
Real clinical context for grounding
🧑‍🔬
Multidisciplinary scientific expertise
Biology, chemistry, clinical — in the feedback loop
📡
Ability to generate new signal
Every assay, every experiment improves the system

Open models are the foundation. Xaira's scientific intelligence is the moat.

X-Cell and X-Scientist are highlighted with the Xaira accent. These are the core proprietary assets that make the architecture uniquely Xaira's. None of this is replicable by an AI lab releasing a general model.

Concrete first step

First pilot: target-validation
intelligence

One system. One workflow. Connected directly to Xaira's core drug-discovery process.

Why this pilot
Connects directly to core drug-discovery process
Demonstrates X-Cell integration
Creates measurable feedback loop
Clear decision utility for scientists
System capabilities
1
Ingest evidence
Internal and external evidence ingestion
2
Reason over X-Cell
Perturbation data and causal virtual cell
3
Generate hypotheses
Competing mechanistic hypotheses
4
Identify uncertainties
Critical unknowns that block decisions
5
Recommend experiments
Discriminating assays to resolve uncertainty
6
Coordinate and interpret
Agents + tools · Interpret results back into decision
7
Support the decision
Advance · modify · stop — with evidence
The target-validation pilot is not a demo — it is a real system connected to a real decision. The goal is the closed-loop learning system, not a proof-of-concept presentation.

Execution plan

Six-month roadmap

Foundation
  • · Model evaluation suite design
  • · Secure infrastructure setup
  • · Team designation
  • · License and legal review for candidate models
Build + Test
  • · Target-validation intelligence pilot: build
  • · X-Cell integration
  • · Internal test with scientific team
  • · Iteration based on feedback
Prospective Pilot
  • · Deploy on active therapeutic program
  • · Prospective evaluation against outcomes
  • · First feedback loop measurement
  • · Report back to leadership: scale, redirect, or stop
Start Month 2 Month 4 Month 6 · Decision point
Six months is the minimum viable test. Month 6 is an explicit decision point — we are asking leadership to commit to six months, not to a permanent program. Scale, redirect, or stop.

The ask

We are asking for six months
to prove this works.

Approve the Open Frontier Intelligence initiative
Appoint an executive sponsor
Designate technical and scientific owner
Allocate a small senior cross-functional team
Establish secure model-evaluation infrastructure
Select the target-validation intelligence workflow as the flagship pilot

"We do not need to own general intelligence. We need to own what general intelligence becomes inside Xaira."

Decision at 6 months

Return to leadership with evidence.
Scale the program · redirect the focus · or stop.
This is a deliberate, bounded commitment.

This is the close. Be explicit: we are asking for six months and a small team. We commit to a clear evaluation at the end. Scale, redirect, or stop — leadership retains the decision.

Appendix A · Verified July 22, 2026

Detailed model specifications

Model Developer Country Release Weights License Commercial Context Architecture Modalities
DeepSeek-V4-Pro DeepSeek China 2026 Yes DeepSeek License Review 1M tokens MoE 1.6T/49B active · CSA+HCA · 32T tokens trained Text
Qwen3-235B-A22B Alibaba / Qwen China 2025 Yes Apache 2.0 Yes 32K native / 131K YaRN MoE 235B/22B active · 128 experts · 94 layers Text
GLM-5.2 Z.ai (Zhipu AI) China Jun 13 2026 Yes MIT Yes 1M tokens (128K out) MoE · IndexShare sparse attn · MTP Text
Kimi K2.6 Moonshot AI China 2026 Yes Custom · Review Review 256K tokens MoE 1T/32B active · 61 layers · MoonViT 400M Text · Vision
Llama 4 Maverick Meta USA Apr 5 2025 Yes Llama 4 Community Yes 1M tokens MoE 17B active/400B total · 128 experts · early fusion Text · Image
Inkling Thinking Machines Lab USA Jul 2026 Yes Custom AUP Review 1M tokens MoE 975B/41B active · 256 experts · 66 layers · hybrid local/global attn Text · Image · Audio

Closed frontier models (API only — no weights)

ModelDeveloperWeightsNotes
GPT-5 / o3 OpenAI None Frontier reasoning · API only · data governance review required
Claude Sonnet 4 Anthropic None Strong safety + reasoning · API only · data governance review required
Gemini 2.5 Pro Google None Multimodal frontier · API only · data governance review required

* Benchmark figures from model cards are developer self-reported and not independently verified. All license notes advisory only — legal review required before deployment.

Appendix. Full verified specs. Kimi K3 excluded — not released as of July 22, 2026. All license notes are advisory.

Appendix B · Legal review required

License details per model

Permissive Apache 2.0 — Qwen3

Commercial use allowed. Attribution required. Patent rights granted. No trademark use of "Qwen." Standard Apache terms apply.

Permissive MIT License — GLM-5.2

Commercial use allowed. Minimal restrictions. Retain copyright notice. Broadly permissive.

Commercial OK Llama 4 Community License — Meta

Commercial use allowed. Restrictions apply for services with >700M monthly active users. Acceptable Use Policy must be followed. Available at github.com/meta-llama.

Review Required DeepSeek Model License

Open weights available. License terms require review for commercial/internal deployment. Regulatory and supply-chain considerations apply for models from Chinese developers.

Review Required Custom License — Kimi K2.6

Open weights available but custom license. Commercial and internal use terms require review. Verify on HuggingFace moonshotai/Kimi-K2.6 before deployment.

Review Required Acceptable Use Policy — Inkling

Custom AUP from Thinking Machines Lab. NOT MIT or Apache. Weights downloadable but commercial deployment requires review of AUP terms. Safety evaluations conducted by developer.

All license information is advisory only. Consult legal before any deployment.

License review is mandatory before any deployment. Green = cleaner path. Warn = requires legal sign-off. Chinese models require additional cybersecurity and regulatory review beyond just the license terms.

Appendix C · If target-validation is not the right first pilot

Alternative pilot ideas

Scientific literature synthesis agent

Deploy a long-context open-weight model to synthesize internal and external scientific publications on a specific disease area or target class. Tests model integration and RAG pipeline without requiring X-Cell coupling.

Entry barrier: low · Decision utility: medium · X-Cell integration: no
Experiment design assistant

Build a system that takes a scientific question and proposes assay designs with justifications. Integrates with internal protocol databases and external literature. Validates against scientist judgment.

Entry barrier: medium · Decision utility: high · X-Cell integration: optional
Multi-model scientific reasoning benchmark

Before any deployment, run a structured evaluation of candidate models on proprietary Xaira scientific tasks. Outputs a ranked comparison informing all future model selection decisions.

Entry barrier: low · Decision utility: high (infrastructure) · X-Cell integration: optional
X-Atlas perturbation reasoning

Direct integration with X-Atlas CRISPR data. Ask open-weight models to reason over perturbation profiles, predict phenotypic consequences, and identify mechanistic hypotheses. Direct X-Cell coupling.

Entry barrier: medium-high · Decision utility: high · X-Cell integration: yes
Alternative pilots in case leadership wants to start with something lower risk before the full target-validation system. The multi-model benchmark is particularly good as a first step since it generates infrastructure value regardless of which pilot is chosen.

Appendix D · Verified July 22, 2026

Sources and verification notes

GLM-5.2
  • HuggingFace: zai-org/GLM-5.2
  • z.ai blog (release announcement)
  • arXiv: 2602.15763
  • License: MIT (verified on HuggingFace)
  • Release: June 13, 2026 (verified)
Kimi K2.6
  • HuggingFace: moonshotai/Kimi-K2.6
  • Architecture: MoE 1T/32B active, 61 layers
  • MoonViT: 400M params
  • Kimi K3 status: announced, not released as of July 22, 2026
Inkling (Thinking Machines Lab)
  • thinkingmachines.ai/inkling
  • HuggingFace: thinkingmachines/Inkling
  • Updated ~July 20, 2026
  • License: Custom Acceptable Use Policy (NOT MIT)
  • Note: correct name is "Inkling" not "Inking"
DeepSeek-V4-Pro
  • HuggingFace: deepseek-ai/DeepSeek-V4-Pro
  • arXiv: 2606.19348
  • Training: 32T tokens, GRPO post-training
  • License: DeepSeek Model License (review required)
Qwen3-235B-A22B
  • HuggingFace: Qwen/Qwen3-235B-A22B
  • qwenlm.github.io/blog/qwen3
  • License: Apache 2.0 (verified from LICENSE file)
  • Context: 32K native, 131K with YaRN
Llama 4 Maverick
  • HuggingFace: meta-llama/Llama-4-Maverick-17B-128E-Instruct
  • Release: April 5, 2025
  • License: Llama 4 Community License Agreement
  • Context: 1M tokens (Maverick), 10M (Scout)
All benchmark figures cited from model cards are developer self-reported. Web search unavailable during research — verified from HuggingFace model pages and official developer sites directly. Legal review of all licenses required before deployment.
Source appendix. All key claims traced to primary sources. No secondary sources used where avoidable. Kimi K3 excluded due to no public weights.