AI Radar for Companies — Monday, September 7, 2026
noticias radar ia modelos de ia ciberseguridad automatizacion

AI Radar for Companies — Monday, September 7, 2026

· CompaniesAutomation

OpenAI's 99.9% score for GPT-6 Astra on ARC-AGI-3 came from a proprietary adapter: with a neutral setup, the same model drops to 62.7%, and independent evaluators can't even agree with each other. Meta is offering a 95% discount on Muse Spark to those who let them train on their prompts and responses. Fluidstack closes a $1.5 billion round led by Jane Street, doubling its valuation to $18 billion without owning a single chip. DeepSeek orders 160,000 Huawei accelerators for a one-gigawatt center in Inner Mongolia. And on Friday, the 24-hour clock for notifying exploited vulnerabilities starts across the EU, with fines up to 15 million euros.

The Sunday radar was about what can be verified; today's is about the fine print. Three threads open the week. One: the 99.9% OpenAI headlined for its new model launch came from a testing setup it doesn't sell, and with the neutral standard, the same model sits at 62.7%. Two: cheap AI is no longer paid for just with money—Meta offers a 95% discount to those who gift their prompts, and DeepSeek orders 160,000 Chinese chips so the inference you buy runs in Inner Mongolia. And three, this Friday a twenty-four-hour clock starts across the European Union for notifying exploited vulnerabilities, with fines of up to fifteen million euros. In money news, another infrastructure round: 1.5 billion dollars for a company that doesn't own a single chip.

62.7% vs. 99.9%: The Number You Buy Isn't the One You Were Shown

OpenAI introduced GPT-6 Astra with a 99.9% score on ARC-AGI-3, the general reasoning test the industry uses as an AGI thermometer. The independent review published on Friday clarifies where that figure comes from: it was achieved with a proprietary adapter that maintains reasoning state between actions, something the neutral setup used to measure rival models does not do. With that standard setup, Astra drops to 62.7%. And independent evaluators don't even agree with each other: Epoch AI, which aggregates over fifty tests, puts it first with 169 points compared to Claude Fable 5.1's 163; Artificial Analysis places it second, 61 against 66, and also behind in its programming agent index, 67 against 70. The model also costs two and a half times more per unit of text than the previous generation. For your company: no manufacturer benchmark is your benchmark. What decides if you switch models isn't a lab percentage, it's your process with your data: take twenty closed real cases from your operations—twenty invoices, twenty orders, twenty tickets—write the correct response for each, and run them through both candidate models measuring accuracy, cost per case, and seconds. It's half a day's work and it's the only number you can defend in a committee. If the salesperson brings you a figure, ask for the neutral setup one and the cost per completed task, not per token. Source.

Meta Discounts 95% If You Let Them Keep Your Prompts and Responses

Meta has opened a "contributor" tier for Muse Spark, its agentic programming model: the price per million input tokens drops from $1.25 to $0.10 and output from $4.25 to $0.20, a twenty-one-fold cut on the expensive part. The condition is written unambiguously: in this tier, prompts, responses, and usage patterns enter the training pipeline for future models. The reason is transparent—Facebook, Instagram, and WhatsApp give Meta plenty of conversation and photos, but almost nothing of what a programming agent needs to improve: real sessions of people fixing real code. For your company: this isn't decided with a calculator, it's decided with a two-line written policy. Divide your work into two boxes: what can leave (internal scripts, tests, disposable prototypes, back-office automations) and what cannot leave under any circumstances (client code under NDA, anything with personal data, and everything that is your competitive edge). The first box can be moved to the cheap tier for real savings; the second cannot, even if the discount is ninety-five percent. And check if your client contracts even allow you to make that decision: many NDAs prohibit transferring code derivatives, and the discount becomes expensive when you pay it in damages. Source.

1.5 Billion for a Company Without Chips: Valuation Doubles in Two Months

Fluidstack closed a $1.5 billion round on Thursday led by Jane Street, valuing it at over $18 billion. In July, it had raised $750 million at a $7.5 billion valuation: it has more than doubled in two months, totaling over $2.6 billion raised. The interesting part is the model: the company doesn't own the accelerators, it builds and operates high-performance clusters for others, and its revenue has gone from $1.8 million in 2022 to $66.2 million in 2024, with a projected $660 million this year. Behind this is a multi-year, $50 billion capacity agreement signed with Anthropic to build custom centers in Texas and New York. The detail to remember is who is putting up the money: Jane Street, a trading firm, has already committed about $6 billion to CoreWeave and $13 billion to Crusoe. For your company: the money financing the compute you consume no longer comes from manufacturers or telecom operators; it comes from financial desks that buy capacity like any other asset and expect a scheduled return. That's neither good nor bad in itself, but it explains why promotions have expiration dates and why it's wise to budget for 2027 using list prices. And it provides a specific question for your next renewal: in which cloud and in which region is your inference actually running, and how much written notice do you get if the price changes. Source.

DeepSeek Orders 160,000 Huawei Chips: Where Your Inference Runs Is a Clause, Not a Detail

DeepSeek is preparing to deploy at least 160,000 Huawei Ascend 950DT accelerators in a roughly one-gigawatt data center in Ulanqab, Inner Mongolia, Bloomberg reported Friday. This would be the largest known cluster built on Huawei silicon, far exceeding previous installations measured in tens of thousands of cards. Two nuances matter: the 950DT was designed by Huawei for training, but DeepSeek wants it to serve its already trained models, and memory shortages will limit production, so the order may take over a year to complete—the company expects to have part of it running by late 2027 or early 2028. For your company: many small businesses have plugged in Chinese models based on price, directly or through an aggregator, without looking at where the call is executed. If that inference ends up running in Chinese territory, the conversation stops being technical: it becomes an international data transfer that GDPR requires you to document, and to which the law of the country where the machine is located applies. Two tasks for this week: ask every provider in writing for the execution region and retention policy, and have a European alternative for the same model ready—open weights can be served from the EU—for processes involving personal data or trade secrets. Source.

The 24-Hour Clock of the Cyber Resilience Act Starts Friday

On September 11, the notification obligations of the European Cyber Resilience Act come into effect. From that day on, any manufacturer of a product with digital elements sold in the Union must notify ENISA and their national incident response team within twenty-four hours of becoming aware of an actively exploited vulnerability or a serious incident; full notification is due at seventy-two hours and the final report fourteen days after a patch is available. Fines reach fifteen million euros or 2.5% of global turnover. Two practical warnings: the deadline counts from when you find out, not from when ENISA's single notification platform is operational—it wasn't at the end of August and launches the same day—and "manufacturer" includes many SMEs that don't see themselves as such: if you sell a product with firmware, your own app, or software packaged under your brand, you are one. For your company: if you manufacture, three things before Friday: identify your reference CSIRT based on your main establishment, appoint a person and a backup who can notify within twenty-four hours even in August, and have the early warning template pre-written, because writing it with the clock running is how those first hours are lost. And if you only buy, prepare for the other side: starting Friday, you will receive notices from your providers that may trigger your own obligations under NIS2, so decide now who reads them and which inbox they go into. Source.

What to Watch This Week?

Whether Anthropic or Google publish their ARC-AGI-3 figures with the neutral setup, which is the only way to truly compare. And on Friday, how many notifications actually come through the ENISA platform in its first hours: it will be the first indication of whether the twenty-four-hour clock is a real obligation or just another box to check.

Watch the 1-minute video