Andrew's Notes
Search
Search
Dark mode
Light mode
Explorer
benchmarks
4 items with this tag.
Aug 24, 2026
AI Agents That Matter
ai-agents
evaluations
benchmarks
llm
reproducibility
overfitting
cost-analysis
princeton
Aug 24, 2026
This is the most misunderstood graph in AI
metr
time-horizon
evaluations
ai-progress
science-communication
benchmarks
mit-technology-review
Aug 24, 2026
Time Horizon 1.1
metr
evaluations
time-horizon
benchmarks
ai-progress
methodology
hcast
Jul 29, 2026
How do we prevent AI agents from going rogue? It starts with a new kind of measurement
ai
ai-safety
ai-alignment
openai
cybersecurity
benchmarks
opinion