The Oct 8 AI briefing rounds up the past 24 hours: cheaper models and on-device agents advanced at the same time. Anthropic released a lower-priced Haiku 5.5 and added API credits for subscribers, NVIDIA and Microsoft built an agent runtime into Windows PCs, and a report on AI penetration tools hitting South Korean banks is a reminder that defenses must keep pace. All seven items below were published or reported on Oct 7-8, 2026 and can be traced to their original sources; stories already covered as standalone pieces today, such as the GPT-6 chat release and Perplexity's embedding models, are not repeated.
Anthropic releases Claude Haiku 5.5: 1M context at a new low price
Anthropic shipped Claude Haiku 5.5 on Oct 7 via the Claude Code v2.1.293 GitHub release, making it the default Haiku model on the Anthropic API: a 1M-token context, priced at $0.10 per million input tokens and $0.50 per million output tokens, rising to $0.50 and $2.50 when prompts exceed 100K tokens. The same day, official accounts announced two cost changes: cache reads on Claude Sonnet 5.5 were halved to $0.10 per million tokens, which the company says cuts long-running task costs by roughly 20%, and Claude Max and Team plans gained monthly Platform API credits - $100 for Max 5x, $200 for Max 20x, and up to $500 shareable on Team - usable with any model including Haiku 5.5. High-volume teams get direct savings; individual subscribers effectively gain extra programmable quota.
NVIDIA and Microsoft launch RTX Spark: Windows PCs rebuilt for agents
NVIDIA announced on its official blog on Oct 7 that it is co-building Windows agent PCs with Microsoft. The new RTX Spark platform pairs a Blackwell RTX GPU with up to 6,144 cores and an up to 20-core Grace CPU, with up to 128GB of unified memory and about one petaflop of FP4 compute. Laptops are open for preorder and ship Oct 16, compact desktops follow in November, with machines from Acer, ASUS, Dell, HP, Lenovo, Microsoft, MSI and Gigabyte. On the software side, Microsoft Execution Containers (MXC) reached general availability, letting agents run persistently in the background under OS control. A Windows DGX Station was also previewed, with a GB300 chip, 748GB of coherent memory and up to 20 petaFLOPS, enough for trillion-parameter models locally. The agent battleground is moving from cloud to desktop.
Report: a lone attacker used AI penetration tools against South Korean banks
The Decoder reported on Oct 8, citing CrowdStrike, that between late September and early October 2026 a suspected single attacker used ARTEX, an open-source AI penetration-testing tool, against several South Korean financial institutions and stole large volumes of data; at Shinhan Bank alone, more than 25,000 records with names, contacts, income and credit limits were taken. The financial regulator held an emergency meeting and the president called for a full investigation. The report argues such tools can find vulnerabilities automatically, letting one person do in days what once took a team. Details remain based on the security firm and media reporting, with full confirmation from affected institutions still pending, but AI tools have clearly changed the tempo of financial cyberattacks.
Fine-tuned Nemotron systems reach gold level at both IOI and IMO
NVIDIA's Nemotron team wrote on Hugging Face on Oct 7 that systems built from Nemotron 3 with supervised fine-tuning, reinforcement learning and feedback-driven inference scored 535.4 out of 600 at IOI 2026, above the gold threshold and the top human score in an unofficial live run, and 30 out of 42 at IMO 2026, above the official gold threshold of 29, with proofs graded by official IMO graders. Models, data and recipes are open on Hugging Face. For the open ecosystem the lesson is clear: with domain data plus a generate-verify-refine loop at inference time, a general model can reach elite competition level without retraining the base.
Microsoft Research open-sources Agent Lightning v1.0
Microsoft Research Asia published Agent Lightning v1.0 on its official research blog on Oct 7, at roughly 3,500 lines of code. It proposes training agents with reinforcement learning using the very harness that runs them in deployment, so developers do not rewrite agent logic inside a training framework. That lowers the engineering cost of continuously training real business agents and suits teams that already ship agent products and want to improve them with real interaction data. The framework is open source and deliberately lightweight.
Google opens the SynthID Detector portal for images, video and audio
Google explained on Oct 7 via its official account that the SynthID Detector portal is open: users upload an image, video or audio file and the portal scans for SynthID watermarks from Google or its partners to indicate whether content was AI-generated. It checks watermarks, not truth itself - no watermark does not prove human authorship - but for editors, moderation and brand teams needing a quick first screen of asset origins, it is a ready official entry point.
GPT-6 Luna Decisions lands on OpenRouter
OpenRouter announced on Oct 7 that GPT-6 Luna Decisions is live: a decision interface for applications that takes text, JSON or images and returns typed choices with probabilities, for picking among models, tools or actions. Pricing is $0.10 per million input tokens with free output and a 1M context; OpenRouter, citing OpenAI's developer account, says decisions can be up to 10 times faster than calling GPT-6 Luna through the Responses API. Apps doing heavy routing and classification should test latency and accuracy on small traffic before switching.