The Hardest Unsolved Problems in Meta, Llama, and Open-Weight AI
Publishing a model's weights is a one-way action: it cannot be recalled, only reacted to. Four problems in that fact remain unsolved, and none of them are solved by a better license.
Publishing a model's weights is a one-way action: it cannot be recalled, only reacted to. Four problems in that fact remain unsolved, and none of them are solved by a better license.
A benchmark score looks like a measurement. Comparing two frontier models honestly requires a shared unit, a clean test set, and a matched operating point, and no laboratory has all three yet.
Fossils, ancient DNA, and stone tools each reconstruct human origins from a different kind of evidence, with different reach and different blind spots — and none of the three can replace the other two.
Probabilistic tournaments, narrative scenario planning, and prediction markets all claim to prepare institutions for an uncertain future — on different dimensions, with different failure modes, and none as a replacement for the others.
Field decades, lab generations, and genomes read backward: three ways of catching selection at work, and what each one can and cannot tell you.
How zooarchaeologists, ancient-DNA labs, and settlement-survey teams actually produce the numbers behind "domestication," "urban population," and "collapse"—method by method, with the uncertainty left in.
Three grounded bets about how the archaeology of early complexity will move by 2035 — on remote sensing, revisionist theories of hierarchy, and the ancient-DNA sampling gap — each with a horizon, assumptions, and a stated way to be proven wrong.
Agent benchmarks report dramatic swings in success rates from one model generation to the next. Look closely, and the harness, the trial count, and what counts as a pass turn out to matter as much as the model does.
Knockouts, CRISPR screens, and GWAS answer different questions about a gene. None of them replaces the others, and mixing up what each one can prove is where a lot of bad biology gets published.
Cultural-evolution theory, cognitive-offloading experiments, and futures scenario scholarship all study how humans and AI are reshaping each other — on different timescales, with different evidence, and different jobs to do.