Andrew's Notes

benchmarks

4 items with this tag.

  • Aug 24, 2026

    AI Agents That Matter

    • ai-agents
    • evaluations
    • benchmarks
    • llm
    • reproducibility
    • overfitting
    • cost-analysis
    • princeton
  • Aug 24, 2026

    This is the most misunderstood graph in AI

    • metr
    • time-horizon
    • evaluations
    • ai-progress
    • science-communication
    • benchmarks
    • mit-technology-review
  • Aug 24, 2026

    Time Horizon 1.1

    • metr
    • evaluations
    • time-horizon
    • benchmarks
    • ai-progress
    • methodology
    • hcast
  • Jul 29, 2026

    How do we prevent AI agents from going rogue? It starts with a new kind of measurement

    • ai
    • ai-safety
    • ai-alignment
    • openai
    • cybersecurity
    • benchmarks
    • opinion

Created with Quartz v5.0.0 © 2026