Project Beacon was announced by Baseten on its official blog on October 9, 2026, in partnership with Goodfire AI, with the goal of building proactive safety controls directly into open-model inference. Instead of adding a text filter after the output, it reads the model's internal activations during generation, using probes to spot a risky event as it happens and act before the content reaches a user or a tool.
The four risks it watches
Baseten lists inference-layer risks including prompt injection hidden in documents or tool results that tries to redirect an agent; an agent proposing actions beyond what its user or organization authorized; customer or company sensitive information appearing where it should not; and early signs of cyber misuse in security-sensitive workflows. Each risk calls for a different response, so Beacon connects the behavior-specific monitors developed by Goodfire to the systems at Baseten that decide how an application responds: request human approval, refuse, fall back to another option, or simply log the event for admin review.
Why monitoring runs in parallel with generation
In the architecture, activation-based and text monitors classify each request at the same time, and a frontier model weighs in only when the two disagree; monitoring runs in parallel with generation and does not block the response. For enterprise customers, Baseten plans a centrally applied safety baseline, policies adjustable per workload, and a clear record of what was flagged and how it was handled; for developers, safety events and controls will arrive through the same APIs and tools used for inference. The example given is concrete: a payment-dispute agent reads merchant messages and account records, malicious content could redirect it, and a wrong action could expose data or move money — reviewing every step with another large model would not scale on cost or latency.
Over the next several months, Beacon's capabilities will start with selected models and monitored behaviors, expand into enterprise controls and developer-facing experiences, and reach the market first through a small number of early partners. For teams building agents on open models, the signal in this update is that safety checks are moving from bolt-on filters into inference infrastructure itself — and whether monitoring signals can be called directly by your own policies will sit alongside price and speed when inference platforms are compared.