Architectures & training
New model architectures, training methods, and scaling experiments at the edge of what small teams can run.
Collaborative AI research lab · est. 2017
Alphabell coordinates hundreds of independent researchers, data scientists, and hackers across dozens of countries. We bet the next AI breakthroughs will come from sharper, more elegant ideas — not from building another data centre.
The dominant paradigm in AI today is straightforward: more data, more parameters, more GPUs. It has carried the field a long way. It is also hiding a quieter truth — that "we don't yet have a good idea" keeps getting mistaken for "we don't yet have enough compute".
We think the next round of real breakthroughs will come from people with sharper hypotheses, careful theory, and the kind of focused thinking that doesn't get easier just because you spent another billion dollars. The bottleneck is the idea, not the cluster.
So we don't try to build a frontier lab. We coordinate as many great thinkers as we can find — independent researchers, data scientists, and hackers across dozens of countries — instead of pouring billions into another data centre.
Read the full argument →The next AI breakthroughs will come from sharper ideas, not bigger clusters. Many minds beat more compute.
The work we like best is the kind that doesn't get cheaper at scale: a clean hypothesis, a careful replication, an interpretability result that finally explains something, a benchmark that exposes a real failure mode.
None of these get unblocked by 10× the GPU budget. They get unblocked when you find the right person and give them runway to think.
That is the entire business model. Find the people. Give them runway. Publish what they find.
We pick problems where the bottleneck is the hypothesis, not the cluster — open questions where the cost of being wrong is low, the cost of being right is high, and the work benefits from being done in public.
New model architectures, training methods, and scaling experiments at the edge of what small teams can run.
Sparse-feature atlases, causal scrubbing, refusal direction studies. Open dictionaries for open models.
Held-out benchmarks, agent-trajectory evals, multilingual prompt sets, blind-spot detection for VLMs.
Long-horizon agent training, schema-aware tool use, replay buffers, and harnesses that don't fall apart on day five.
New training data, new ways to study existing data, and underserved-language eval sets built with native speakers.
End-to-end replications of published work, bug bounties on shipping code, and shared reference numbers.
A sample of what members have released — datasets, audits, atlases, tools. Every release is reproducible by design: code, seeds, configs, and held-out splits all live next to the paper.
An open atlas of 18,000+ dictionary features across four popular open-weight 7B checkpoints. CC-BY release.
End-to-end replications of nine recent router-distillation papers. Three reproduce cleanly, two partially, one fails with the published code. All notebooks public.
Tracking the lifetime of dictionary features across pretraining checkpoints. Code, mid-training snapshots, and intermediate analyses released monthly.
18,000 labelled adversarial tool-use prompts in Cantonese — the first public dataset of its kind. Dual-released to Hugging Face and OpenReview.
Training a 7B agent on a 14-day partial-observability trajectory replay buffer. Eval harness and seed trajectories already public; checkpoints due Q3.
Yoruba, Igbo, Hausa, Swahili, and Amharic code-switching prompts for agent tool-use. Built from real customer-service transcripts with verified native-speaker review.
We take reproducibility as a design constraint, not a virtue badge to be added at the end. Every project releases enough material for an independent team to verify the result — and we fund the people who try.
Code, weights, datasets, and write-ups all release under permissive licenses. Closed-source output is the exception, and we want a reason on the record.
No "and we did some hyperparameter tuning". Every release ships the seed, the config, the data split, and the version of the eval harness used to score it.
We pay members to replicate published work — ours and others'. Negative results are published the same way positive ones are. Outcomes go in the public log either way.
Every project releases enough material for an independent team to verify the result. Reproducibility is treated as a first-class output, not a footnote.
Membership has no formal threshold. A high-school student writing a clean implementation, a PhD on sabbatical, and a senior engineer chipping in evenings all participate on equal footing — judged by the work.
Time and runway for one idea.
Short, low-paperwork support for a focused project — a benchmark, a tool, a replication, a paper. Single-page application, decisions in weeks rather than months.
How they workFunded time to commit deeper.
Three to twelve months of supported research time, with mentorship and compute. For people who want to dig in on an Alphabell project or pursue their own agenda.
Read moreHead-to-head, in the open.
Recurring sprints with public leaderboards and held-out evals. Strong runs become a credential — and the submissions themselves become reusable artefacts that other researchers can build on.
Active boardsThe biggest barrier, lifted.
GPU clusters, evaluation harnesses, datasets, and engineering support — available to members with active projects. The single largest practical barrier facing independent AI researchers.
How it works