Computer Science + Economics · Illinois
We use language models and agents to measure what AI can do, what it costs and who is using it, then forecast where it goes next. We also build tools that make AI cheaper and easier to adopt.
Vision
Anyone should be able to see where AI is heading, and afford to use it.
Mapping the research literature on self-driving laboratories: citation and keyword expansion from curated seed lists, LLM classification at scale, and full-text retrieval. Round one found about 1,000 relevant papers; round two tests whether more are still out there.
Career outcomes of computer science PhDs from the top 100 US departments, broken down by research area, built by linking dissertations to career records.
Every paper and blog post from 19 LLM developers in one searchable timeline, each linked to its code and open weights, to see who releases what and when.
When do open-weight models catch the closed labs? Twenty months of a public leaderboard rebuilt from archived snapshots, every model scored on one fixed benchmark at a time. Interactive page
Our own capability scale, fitted with item response theory on independently measured benchmark scores, so catch-up times do not depend on any one leaderboard's weighting. Validation thresholds are registered before the checks run.
A theorem-proving benchmark cut from a frontier-mathematics Lean formalization published after every open-weight checkpoint we test. 504 problems, graded by the Lean compiler with no LLM judge.
Same agent, same model, same 20 file and text tasks, on Linux, macOS and Windows. Whatever differs in the results is the operating system.
Training-data contamination tests across 10 open-weight base models, and seven open-source OCR pipelines compared on 100 journal articles.
In real agent coding sessions, serving cost is dominated by prefill (roughly 300 to 400 input tokens per output token) and repeated fixed prefixes, a profile unlike chat traffic.
How much of frontier AI a lab can run on its own modest, older hardware: a 744B mixture-of-experts language model at 1-bit, 262K-token context on a 27B model, and 33B video generation models without quantization. What works, what breaks, and what it costs.
The token cost of wrapping the same information in XML, JSON, OOXML or HTML before it reaches a model.
A short film of Quad Day, the start-of-year fair on the Illinois main quad, generated end to end with open video models on our own GPUs. A test of how far a small team can get with open tools and a script.
A 450 by 600 metre stretch of the Illinois main quad as a walkable 3D gaussian splat, built only from a public airborne lidar survey, with no street-level photos. Walk it
How far a regional 3D model can get on public data alone, and where it falls short, to decide whether collecting street-level imagery is worth it.
A browser workspace with a coding agent for every student, running on shared open models. Built around the data chores economics students actually do, like cleaning a government statistics table.
Connecting everyday agent tools to the university's free open-model service. We added an Anthropic Messages endpoint to Lumen and package Claude Code to run on it. Code
Real-time transcription and translation for English talks and meetings, running on a local GPU, with a searchable transcript afterwards.
Open-source AI teammates that share one cloud computer with a shell, files and a browser each, message each other, run on schedules, and ask before doing anything irreversible.