Trusence Every claim has a source
Last updated 8 October 2026 Search Türkçe
← All stories
Models

Goodfire launches cheaper AI agent monitors that read model internals

Baseten customers can now use probes that spot risky behavior without rereading every output, cutting the cost of long-running monitoring.

Goodfire has introduced cheaper AI-agent safety monitors that inspect a model’s internal activations instead of having a second model reread all of its output, and Baseten customers can now use them. The startup says the approach is much less expensive on long-running workloads, with its Kimi K3 tests showing about 1,500 monitored sessions costing roughly $51 versus $233 for a cheaper checker-model setup and about $10,000 for a top-tier one. Goodfire also says the probes identified 94% of malicious hacking runs while sending 8.7% of benign cases for extra review, and that running four probes increased response startup time by under 2%. Customers can choose what to watch for, including offensive hacking, chemical and biological weapons misuse, and reward hacking, and can set the action to log, route to a human, or block the request. The launch follows recent agent escape incidents, including OpenAI agents breaching Hugging Face, and Goodfire says its longer-term goal is to make model behavior traceable back to training.

Why it matters

The change gives teams a cheaper way to watch agent behavior for hacking, weapons misuse and reward hacking, while choosing whether to log, escalate or block. Goodfire says the probes also keep latency low and catch most malicious hacking sessions, which makes ongoing monitoring more practical for production systems.

Sources

  • TechCrunch