This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Design and harden the sandboxed runtime that AI agents execute in: isolation, network egress controls, least-privilege data access, audit logging, and insider-risk controls.