RESEARCH NOTE 003 / NEGATIVE RESULT
Anatomy prediction meets a stronger control
A model can appear to predict anatomy while mostly capturing how large an animal is or which animals it is related to.
We tested head–body length in 32 mammal species spanning 28 families, using 3,604 published gene profiles. Entire families were withheld during fitting. The primary question was whether the genomic representation added predictive value after accounting for body mass and evolutionary relatedness.
The result
| Input or comparison | Error reduction versus mean |
|---|---|
| Gene profiles | 22.5% |
| Nearest relative | 43.8% |
| Body-mass scaling | 96.3% |
Reductions in family-balanced squared error for log-transformed head–body length. Body mass is an observed phenotype supplied to the comparison model; its result is not a prediction from DNA.
No incremental genomic benefit detected: adding gene profiles to body mass plus relatedness increased error by 0.43%. The uncertainty interval crossed zero, and neither permutation diagnostic supported a beneficial genomic contribution.
The very strong mass comparison is a familiar size relationship, not a new biological law. Measurement independence between mass and length is not established: the source compilations may share studies and include extrapolated values.
What the test actually asks
Can this particular comparative sequence representation explain variation in length that a size-and-relatedness model misses? On this small dataset, it did not. That rules out promoting this predictor as a breakthrough. It does not show that DNA cannot explain anatomy.
The head–body-length outcome came from PanTHERIA, with body mass from the Amniote database. The sequence representation came from published gene trees distributed with RERconverge. This was a local CPU analysis of existing public data; we did not train a whole-genome language model or run new biological experiments.
Checks and limits
We fixed the protocol before computing this test’s model metrics, checked source values and species identities, reconstructed saved metrics in a separate audit, and ran 400 permutation diagnostics. The protocol was frozen locally; it was not externally preregistered. Teat count and forearm length were not tested because coverage was insufficient.
These are exploratory results following earlier trait tests. The sample is small, the traits are compiled species summaries, and the fixed comparative gene trees include evaluation-species sequences. Uncertainty intervals are descriptive, and the permutation fractions are not calibrated phylogenetic significance tests. The code audit verifies implementation; it is not an external replication.
What changes next
The next research unit should be a consistently measured structure across independent evolutionary changes. Size, relatedness, stage, and measurement source need to be explicit competing explanations. Reproducing a published structure–sequence link should precede any claim of a new one.
Negative results sharpen that decision. The useful output here is a benchmark that exposes an easy way to overstate genomic prediction.
Public inputs
- PanTHERIA — head–body length and taxonomic context.
- Amniote database — body-mass comparison.
- RERconverge repository — published example amino-acid gene trees.
Publication policy: these notes share aggregate findings, source attribution, and limitations. Internal implementation and future candidate work are not included. Public summaries alone are insufficient to reproduce every analysis.
Explore the founding mission