The Hardest Unsolved Problems in Meta, Llama, and Open-Weight AI
Publishing a model's weights is a one-way action: it cannot be recalled, only reacted to. Four problems in that fact remain unsolved, and none of them are solved by a better license.
Tagged AI governance · show all articles
Publishing a model's weights is a one-way action: it cannot be recalled, only reacted to. Four problems in that fact remain unsolved, and none of them are solved by a better license.
Cultural-evolution theory, cognitive-offloading experiments, and futures scenario scholarship all study how humans and AI are reshaping each other — on different timescales, with different evidence, and different jobs to do.
Refusal training, content classifiers and red-team programs are not proof a deployed system is safe. They are ten separately documented ways such systems keep failing, each with its own paper trail and its own cause.
A misuse policy that lives in a slide deck is not a safety system. This is a practitioner's account of the layered defenses, enforceable policies, self-red-teaming and post-launch instrumentation that make one hold under real traffic.
Two questions decide how AI alignment and safety practice matures by 2035: whether verification moves beyond behaviour, and whether the field becomes an externally certified discipline or stays self-governed by labs. Four scenarios, each with a falsifier.
Two questions decide whether agent evaluation becomes a certified, externally audited discipline by 2035, or stays a signal vendors mostly produce and grade themselves. Four scenarios, each with a falsifier.
Two questions decide how agent architecture matures as an engineering discipline by 2035: pattern convergence or fragmentation, and certified reliability or empirical best-effort. Four scenarios, each with a falsifier.
Every refusal passes through a trained policy, a classifier, account monitoring, and a capability-gated release process — four thresholds, each trading missed misuse against blocked legitimate use, documented in Anthropic's own safeguards and RSP filings.
Two questions, not one, decide how Anthropic and the Claude line look in 2035: does its safety-training method stay on its current path, and does its scaling policy become everyone's floor or stay its own. Four scenarios follow, each with a falsifier.
Every AI rule must first name something: a model, a role, a use case, a compute number, or a running service. Each choice buys enforceability somewhere and loses it somewhere else, and the losses are where the arguments now are.