Andrew's Notes

anthropic

8 items with this tag.

  • Aug 26, 2026

    What just happened? Pragmatism and Pessimization

    • ai-alignment
    • ai-safety
    • openai
    • deepmind
    • anthropic
    • rlhf
    • integrity
    • epistemics
    • lesswrong
  • Aug 24, 2026

    Congressional oversight letter to Anthropic regarding security incidents

    • ai-policy
    • ai-governance
    • congress
    • anthropic
    • cybersecurity
    • ai-incidents
    • oversight
    • evaluations
  • Aug 22, 2026

    Patterns and problems in emerging multiagent systems

    • ai
    • ai-safety
    • ai-agents
    • multi-agent-systems
    • anthropic
    • evaluations
    • collusion
    • coordination
  • Aug 22, 2026

    Risk Report: August 2026

    • ai
    • ai-safety
    • ai-governance
    • anthropic
    • alignment
    • responsible-scaling
    • evaluations
    • biosecurity
  • Aug 22, 2026

    When AI builds itself

    • ai
    • ai-safety
    • ai-governance
    • anthropic
    • recursive-self-improvement
    • automated-rd
    • ai-policy
    • forecasting
  • Aug 08, 2026

    System Card: Claude Mythos Preview

    • ai-safety
    • system-card
    • anthropic
    • responsible-scaling-policy
    • alignment
    • cybersecurity
    • biosecurity
    • dangerous-capabilities
    • model-welfare
  • Aug 06, 2026

    Toy Models of Superposition

    • ai-safety
    • interpretability
    • mechanistic-interpretability
    • superposition
    • polysemanticity
    • neural-networks
    • anthropic
    • research-paper
  • Jul 29, 2026

    To deliberately pace

    • ai
    • ai-safety
    • ai-governance
    • ai-policy
    • open-source-ai
    • eu-ai-act
    • anthropic
    • openai
    • newsletter

Created with Quartz v5.0.0 © 2026