Technology
Screening technology and LULA benchmarks.
How Om library × library screens generate binding data, how LULA scores Om Accessible Space, and the enrichment and attention results.
what we will do
LULA, powered by massive Om data, scores Om Accessible Space, and identifies molecules that can be sourced and validated to bind your target in weeks.
- 01
Protein sequence
start from your target
- 02
Score the space
LULA scores Om Accessible Space
- 03
Select binders
molecules predicted to bind only your protein
- 04
Validation data back
AS-MS binder and non-binder labels in weeks
how screening augments data
Library × library screens train LULA to pick Accessible Space molecules.
We screen hundreds of millions of molecules with Om library × library screening technology. That binding data trains LULA, which then scores Om Accessible Space and selects molecules predicted to bind only your protein.
- 01
Two libraries
encoded protein targets and small molecules, so every pairing can be traced back
- 02
Pico-scale library × library
every molecule against every protein, simultaneously, in droplets
- 03
Sequence and score
sequencing reads the multiplexed result that trains LULA
LULA-1 & LULA-2
Score molecules from 2 Wallet Credits / 1K scored.
What it is
Om’s small-molecule protein interaction foundation model.
What LULA means
Ligand Underpinned Latent Associations.
How it's trained
2D ligand binding data and protein sequence alone, with no structure in training.
What you get
A binding score you can use to rank molecules directly from sequence and SMILES.
Structure-first systems like BoltzBio's Boltz-2 and Nesso-1 model each protein-ligand pair before scoring. LULA embeds the target once, then scores molecules directly from SMILES and sequence, without crystal structures, docking, or folding.
LULA-1
FastScreens broadly across many targets — the default for low-cost, high-throughput molecule scoring.
LULA-2
High-ResScores about 200 molecules/sec for the hardest targets — protein-protein interfaces, flat pockets — where precision matters most.
625×
lower than BoltzBio API with LULA-2 High-Res at 4 Wallet Credits / 1K scored
~200×
faster than Nesso-1 using LULA-2 throughput
200/sec
LULA-2 scoring throughput for direct molecule ranking
1250×
lower than BoltzBio API with LULA-1 Fast at 2 Wallet Credits / 1K scored
References: BoltzBio public pricing lists $0.025/molecule small-molecule pricing; Nesso-1 Technical Report reports roughly one prediction per second on one GPU. Checked July 30, 2026.
what data is powering LULA?
We screen entire proteomes against hundreds of millions of molecules, to power proteome-wide interaction models like LULA-1 and LULA-2.
The technology
Pico-scale library × library screening: every molecule in one library tested against every protein in another, simultaneously.
The economics
Multiplexed at massive scale instead of one compound at a time — orders of magnitude cheaper than the status quo.
Feeds the next model
Every screen becomes training signal — sharpening every future LULA-1 and LULA-2 run.
300M data points per protein
1,000 proteins a month capacity
~900B data points a quarter
enrichment results, verified
Top-ranked molecules carry the signal.
Across proteome-scale evaluation, LULA-1 recovered 267,324 known binders in the top 1,000 ranked molecules per target versus 4,877.6 expected under random ranking. Family AUROC remains strongest across the same druggable target families.
composition and performance · target-family AUROC
evaluated target-family coverage
Ligase/Synthase
0.829· 151
Oxidoreductase
0.811· 424
Transferase
0.808· 187
Ion channel
0.808· 147
Phosphatase
0.808· 134
Hydrolase
0.805· 351
Kinase
0.800· 707
Transporter/Pump
0.796· 191
Receptor/GPCR
0.793· 959
Protease
0.786· 316
Other
0.750· 1,528
Epigenetic/Chromatin
0.727· 207
Nucleic-acid enzyme
0.716· 137
EF@1000
54.8x
Aggregate enrichment in the top 1,000 ranked molecules per target.
Known binders recovered
267K
Recovered in top-ranked candidate sets across the benchmark.
Random baseline
4.9K
Expected known binders under random ranking.
The donut shows evaluated target-family coverage across 5,439 targets; slice size is target share and color tracks AUROC. The benchmark result is measured by known-binder recovery in the top-ranked candidate set.
fast, then precise where it counts
LULA-2 turns hard targets into real hits.
On these five targets LULA-2 recovers 3.3x to 28.9x more known binders in the top 1000 than LULA-1 — a mean enrichment gain of 10.3x.
Top-1000 precision on a 12-target product panel, fixed LULA-2 checkpoint; muted bar is LULA-1, accent bar is LULA-2. LULA-1 covers every target broadly, and LULA-2 is the high-resolution pass for protein-protein interfaces, flat pockets and other targets where selectivity matters.
attention discovery
Attention finds real contacts across 164 independent structures.
LULA-1 trains on ligand binding data and protein sequence alone. This is what its attention finds anyway, checked against real, recent PDB structures it never trained on.
- 01
Trained on binding data alone
LULA-1 learns from real ligand binding measurements and protein sequence. No 3D structure is part of training, ever.
- 02
Attention ranks the residues
Its cross-attention between ligand and protein surfaces which residues correlate most with binding — a sequence-level signal alone.
- 03
Structure models seed a pocket
Those ranked residues, together with the ligand, seed a predicted pocket and binding pose using Chai-1 and other structure models.
- 04
Checked against real structures
The result correlates closely with real, independently solved PDB structures — on targets the model never trained on.












500M training data points
8% of the human proteome
300M molecules screened per protein









