Andrew's Notes

Home

❯

2026-08-24

2026-08-24

8 items under this folder.

  • Aug 24, 2026

    AI Agents That Matter

    • ai-agents
    • evaluations
    • benchmarks
    • llm
    • reproducibility
    • overfitting
    • cost-analysis
    • princeton
  • Aug 24, 2026

    Can AI agents conduct open-ended AI research? Early evidence from two case studies

    • ai-agents
    • ai-research-automation
    • evaluations
    • recursive-self-improvement
    • neurips
    • peer-review
    • crux
    • failure-modes
  • Aug 24, 2026

    Congressional oversight letter to Anthropic regarding security incidents

    • ai-policy
    • ai-governance
    • congress
    • anthropic
    • cybersecurity
    • ai-incidents
    • oversight
    • evaluations
  • Aug 24, 2026

    Incident report: unsanctioned agent behaviour during cyber testing

    • aisi
    • ai-incidents
    • cyber-evaluations
    • ai-agents
    • ai-safety
    • containment
    • social-engineering
    • evaluations
  • Aug 24, 2026

    METR's predeployment evaluation of GPT-5.6 Sol

    • metr
    • evaluations
    • time-horizon
    • openai
    • reward-hacking
    • predeployment-testing
    • ai-safety
  • Aug 24, 2026

    This is the most misunderstood graph in AI

    • metr
    • time-horizon
    • evaluations
    • ai-progress
    • science-communication
    • benchmarks
    • mit-technology-review
  • Aug 24, 2026

    Time Horizon 1.1

    • metr
    • evaluations
    • time-horizon
    • benchmarks
    • ai-progress
    • methodology
    • hcast
  • Aug 24, 2026

    We need a Science of Evals

    • evaluations
    • ai-safety
    • ai-governance
    • methodology
    • capability-elicitation
    • standards
    • apollo-research

Created with Quartz v5.0.0 © 2026