All-Atom Structural Knowledge for Omni Biomolecular Design

Preserve the unified generator. Add pretrained all-atom structural knowledge.

AAA†    Han Liu†
Center for Foundation Models and Generative AI · Northwestern University

A lightweight way to transfer rich all-atom structural representations into a frozen omni biomolecular generator while preserving its native geometric structure.

Target and constraints feed a frozen PairFormer above and the frozen AnewOmni encoder below. Trainable projection MLPs and cross-attention enrich native scalar conditioning. Native latents pass unchanged into frozen AnewOmni diffusion, followed by the frozen decoder and generated structure.
Overview. The visible target feeds two frozen branches: PairFormer provides single and pair structural knowledge, while the AnewOmni encoder supplies native block latents. Only the projection MLPs and cross-attention interface are trained. The schematic groups native condition preparation with the encoder; scalar queries come from native position and block-type conditioning. The frozen diffusion model and decoder retain the geometric generation path. Download diagram ↗

About

Omni biomolecular models aim to support peptides, small molecules, antibodies, and other molecular classes through one common generative interface. AnewOmni does this with a shared block-level latent state. The tradeoff is that fine-grained all-atom structure is compressed before diffusion.

Core question. Can we keep the unified omni generator while adding structural knowledge learned by a stronger all-atom model?
Keep the omni latent

The shared block-level state remains the generative space.

Add all-atom knowledge

Frozen PairFormer features contribute token-level and relational structural information.

Preserve native geometry

The transfer path only changes scalar conditioning; coordinate updates remain native.

Geometry-compatible representation transfer

PairFormer receives the same visible design condition as the generator and produces final-layer single and pair representations. These representations are computed once per design task, aligned to the corresponding AnewOmni blocks, and reused throughout sampling.

Frozen donor

PairFormer

Extract single-token and pairwise all-atom structural representations.

→
Trainable interface

Cross-attention

Native scalar conditions query the donor knowledge. Pair features shape both attention and messages.

→
Frozen generator

AnewOmni

Original latent diffusion model still predicts feature noise and all 3D coordinate updates.

Why this preserves geometry

The transferred information is treated as fixed scalar side information. Rotation acts on AnewOmni’s native coordinates, not on the donor knowledge. Because the interface does not introduce a geometric direction, the augmented coordinate prediction keeps the native transformation law.

\epsilon_{\mathrm{aug},q}\!\left(Rz_t^G,t\mid Rz^C,s,p\right)=R\,\epsilon_{\mathrm{aug},q}\!\left(z_t^G,t\mid z^C,s,p\right)

Representation guidance

At inference time, the base and augmented predictions can be interpolated with a scalar guidance parameter γ.

\epsilon_\gamma=\epsilon_{\mathrm{base}}+\gamma\left(\epsilon_{\mathrm{aug}}-\epsilon_{\mathrm{base}}\right)
Guidance strength
γ = 1.00

Use the learned PairFormer-conditioned model at training strength.

Base prediction
All-atom correction
All-atom knowledge
sᵢ, pᵢⱼ
Omni generator
same latent, richer condition

Theory: conditional geometry compatibility

Assumptions. Fix the extracted PairFormer single and pair representations s and p. Native scalar conditioning a is rotation invariant. The native AnewOmni denoiser has an invariant scalar-noise prediction and a rotation-equivariant coordinate-noise prediction. Rotate generated and context coordinates jointly, including native geometric prompts where applicable.

Theorem — scalar transfer preserves native equivariance

\begin{aligned}\ell_{ij}&=\frac{(W_Qa_i)^\top W_Ks_j}{\sqrt{d}}+b_\phi(p_{ij})\\[5pt]\alpha_{ij}&=\operatorname{softmax}_j(\ell_{ij})\\[5pt]m_i&=\sum_j\alpha_{ij}\left(W_Vs_j+W_Pp_{ij}\right)\\[5pt]\widetilde a_i&=a_i+W_Om_i\end{aligned}

The transferred conditioning stays invariant. Consequently the augmented scalar prediction is invariant and the coordinate prediction is equivariant under every R ∈ SO(3).

\begin{aligned}\epsilon_{\mathrm{aug},h}(Rz\mid R\,\mathrm{context},s,p)&=\epsilon_{\mathrm{aug},h}(z\mid\mathrm{context},s,p)\\[6pt]\epsilon_{\mathrm{aug},q}(Rz\mid R\,\mathrm{context},s,p)&=R\,\epsilon_{\mathrm{aug},q}(z\mid\mathrm{context},s,p)\end{aligned}
Proof

With a, s and p fixed under rotation, the attention logits do not change. Softmax weights, projected values and their weighted sums therefore do not change. The residual conditioning ã is invariant. Substituting ã into the unchanged native denoiser invokes its original symmetry law, proving scalar invariance and coordinate equivariance.

Corollary — representation guidance

\epsilon_\gamma=(1-\gamma)\epsilon_{\mathrm{base}}+\gamma\epsilon_{\mathrm{aug}},\qquad\gamma\geq0

For a fixed scalar γ, both coordinate predictions obey the same rotation law. Their linear combination does too, including extrapolation at γ > 1.

Scope. This is conditional on fixed donor representations. It does not establish end-to-end equivariance when the raw input is rotated and PairFormer features are recomputed, nor does it prove empirical improvement.

See the theorem and detailed appendix proof in the paper, or inspect the proof source.

LNR benchmark results

Four validation-selected adapters and a historical AnewOmni reference, evaluated on the same 93 targets with 10 candidates per target. The adapters use the official generation and reconstruction procedure.

Candidate means within each target, then equal weight across targets. Lower is better except AAR. Bold indicates best; underline indicates second best.
ModelAAR ↑Complex RMSD ↓ÅPeptide RMSD ↓ÅIntra-peptide
clash ↓
Interface
clash ↓
dG ↓REUddG ↓REU
AnewOmni · historical6.517%8.5852.9170.987%0.877%9.3546.69
Single-only6.223%8.5882.8700.934%0.822%-0.9936.34
Pair-only6.412%8.5422.8580.921%0.546%-9.0628.28
No pretrained features6.214%8.5892.8870.880%0.879%1.4038.73
Initial evidence. Single + Pair improves all seven point estimates over the no-pretrained-feature control. Pair-only has the lowest interface clash and mean energies. Against the historical AnewOmni reference, Single + Pair improves the structural and mean energy metrics, while AAR is slightly lower.

These are single-seed estimates, not a statistical-significance claim. The AnewOmni row uses historical candidate indices 0–9 and is not noise-paired with the adapters; it is not a paper-exact reproduction. BoltzGen is not included.

Energy distributions and interpretation

dG is the Rosetta interface energy after FastRelax (ref2015); ddG subtracts the corresponding relaxed reference energy. All five groups were scored with the same evaluator. Units are Rosetta energy units (REU), not experimentally measured free energies. Original structures remain the inputs for the five structural metrics.

ModelMedian dG (REU)Median ddG (REU)dG < 0ddG < 0
AnewOmni · historical-18.2417.4789.46%8.60%
Single + Pair-18.5016.5191.18%9.35%
Single-only-17.7217.4789.25%8.39%
Pair-only-18.5117.0990.86%9.89%
No pretrained features-17.9617.5890.00%9.89%

Means are sensitive to large positive outliers; medians and negative-energy fractions provide additional context. Because every group uses the same reference structures, mean dG and mean ddG have identical rankings and are not independent evidence.

Training results and checkpoint selection

Each adapter was trained from scratch on 10,000 sampled records for 10 epochs. The AnewOmni base and PairFormer donor remained frozen. All four correctness gates passed; fixed validation loss selected the checkpoints.

AdapterSelected epochInitial train lossSelected train lossValidation loss
full91.0754921.0614211.291728
single71.0754961.0670281.294466
pair91.0754931.0618671.290897
control71.0754981.0666931.294433
Evaluation protocol and implementation scope

Official API target order, native-length peptide inputs, batch size 16, 100 diffusion steps, initial VAE decode followed by re-encoding and one fixed-topology reconstruction cycle, 10 decoding steps per stage, and FP32 highest precision. The existing 930 candidates per adapter were retained without quality-based resampling. Energy scoring adds no new generation.

The released adapter modifies scalar conditioning for all aligned blocks, including context blocks. The paper PDF is compiled from the latest author-supplied source ZIP; the verified results on this page are maintained separately. The data and protocol links above describe this completed 10-candidate evaluation.

BibTeX

@misc{allatomomni2026, title={All-Atom Structural Knowledge for Omni Biomolecular Design}, author={AAA and Han Liu}, year={2026} }