AI Radar for Companies — Thursday, August 27, 2026
noticias radar ia chips de ia modelos abiertos automatizacion

AI Radar for Companies — Thursday, August 27, 2026

· CompaniesAutomation

Thursday is about who builds your AI's engine and how much they'll sell it to you for. OpenAI published results for Jalapeño, its first custom inference chip built with Broadcom: between 1.5x and 1.9x peak performance over Nvidia's Blackwell systems and up to 3.6x less latency—though without publishing power consumption and for inference only. IBM released Granite 4.2 under Apache 2.0 license—3B, 8B, and 30B parameters, 128k context tokens, and switchable reasoning—to run on your own server. Amazon is closing Mechanical Turk on September 30, twenty-one years and half a million workers later. And DeepSeek signs roughly $7.4 billion at a $74B pre-money valuation, heading for the STAR Market in 2027.

Wednesday's radar was about who already has agents running; Thursday's is about who manufactures the engine and how much they're going to sell it to you for. OpenAI showed off its first proprietary inference chip yesterday and says it performs between 1.5 and 1.9 times better than Nvidia's Blackwell systems, on the same day that Nvidia presented its results. In parallel, IBM released a family of reasoning models under the Apache 2.0 license that fit on your own server, Amazon closed the human labor market that had been labeling AI data for twenty-one years, and DeepSeek is currently signing a round of about $7.4 billion.

OpenAI now manufactures its own chip: inference pricing no longer depends solely on Nvidia

OpenAI published the results yesterday for Jalapeño, its first proprietary inference accelerator, developed with Broadcom for silicon and networking and with Celestica for systems integration, stemming from the October 2025 agreement for 10 gigawatts of custom accelerators. The figures provided by the company: 1.5 to 1.9 times more peak performance work compared to Blackwell generation systems, 1.7 to 3.6 times less end-to-end latency, and 2.1 to 4.1 times faster in ultra-low latency interactive inference. The caveats weigh as much as the figures: these are performance metrics and not consumption—OpenAI did not publish the wattage—, the chip is for inference only and does not touch training, and deployment begins at the end of this year with volume production in 2027. It was announced the same day Nvidia presented results after the market close. For your company: the cost per token you pay is, ultimately, someone else's inference cost. When the world's largest buyer of accelerators starts making their own, and Google and Amazon have already had theirs for years, the medium-term price direction is only one way. Two specific decisions today: do not sign annual fixed-price consumption commitments with a horizon of more than twelve months, because you are locking in the price for a year in which the provider's cost is going to drop; and maintain a model gateway between your application and your token provider, so that switching providers is a configuration change and not a project. Source

IBM gives away reasoning models that fit on your server: "in-house" AI levels up

IBM released Granite 4.2 on Tuesday: three dense models of 3, 8, and 30 billion parameters, under the Apache 2.0 license—they can be downloaded, fine-tuned, and put into production without commercial restriction—, 128,000 native context tokens, extending to 512,000 in the 30-billion version. The novelty isn't the size but two things: a reasoning mode that switches on and off depending on the task, and an agentic reinforcement learning phase for the two large models within real-world software engineering, terminal, and web search environments—meaning they are trained to use tools rather than just answer. They come with cryptographic signatures, ISO certification, and transparency documentation, and are available on Hugging Face, Ollama, LM Studio, GitHub, and watsonx. For your company: this is exactly for what you don't send to an API today due to contract or fear—payrolls, medical records, case files, customer contracts, supplier prices. The 8-billion model runs on a mid-range GPU and the 3-billion one on a modest machine, so the pilot isn't an investment, it's an afternoon. And for AI Act purposes, it plays in your favor: if the model runs on your infrastructure and comes with its transparency documentation, you hold the traceability that might be required of you, rather than depending on a provider to send it. A warning for an honest calculation: open isn't free; the cost moves from the token bill to the server and the person maintaining it, so compare with your actual volume before moving anything. Source

Amazon closes the market where half a million people labeled AI data

Amazon is closing Mechanical Turk on September 30, twenty-one years after its 2005 launch. The platform grew to have more than 500,000 workers performing microtasks—data labeling, audio and video transcription, surveys—and stopped accepting new customers on July 30, the same day as SageMaker Ground Truth and Amazon Augmented AI, AWS's other two human review and labeling services. The data point that explains the closure better than the press release is from 2023: an EPFL study found that between 33% and 46% of workers on the platform were using language models to complete text tasks they were being paid to do by hand. For your company: if you have verification, moderation, transcription, or labeling outsourced to a microtask platform, you have five weeks to move that process, and it's worth using the transition to rethink it entirely instead of just looking for the same service elsewhere. But the underlying reading is more uncomfortable and applies even if you aren't an Amazon customer: if you buy "human review" as quality control for your AI processes—because a customer, an auditor, or the AI Act itself requires it—demand that your provider explains how they guarantee it is human. If they can't, the control you thought you had is one model reviewing another model, and that isn't a control. Source

DeepSeek signs 7.4 billion at a 74 billion valuation and prepares for 2027 IPO

DeepSeek is currently closing its second round: about 50 billion yuan—roughly $7.4 billion—on a pre-money valuation of 500 billion yuan, about $74 billion, with the signing expected before the end of August. This represents a 43% valuation increase since the June round, closed just two months ago. Investors like Monolith and Shixiang Capital funds and battery manufacturer CATL are returning, and talks are underway for CPE, Legend Capital, and Stony Creek Capital—a semiconductor-focused fund—to join. The operation is tied to a schedule: the company is preparing for an IPO on Shanghai's STAR Market as early as the second quarter of 2027. For your company: there is no investment recommendation here; the useful fact is different. The provider that set the industry's price floor has just secured capital for at least a couple more years, with battery and semiconductor money involved, meaning the downward pressure on open model prices won't ease in 2027. It's the same conclusion as the OpenAI chip at the other end of the chain: what seems expensive to you today will cost less, so buy flexibility rather than discounts. And a governance note that is often confused: downloading the weights of a Chinese model and running them on your server in Europe vs calling an API hosted in China are two different things for GDPR and your processing records. The former is a technical decision; the latter is an international data transfer. Source

What to watch tomorrow?

Nvidia presented results last night after the close: what matters isn't the quarter but the guidance for the next one and whether Jensen Huang mentions his largest customers' proprietary silicon without being asked. And a second, quieter one: Reuters reported yesterday that Moonshot is negotiating with Microsoft, Amazon, and Google to host its Kimi K3 model on Azure, AWS, and Google Cloud in exchange for up to 30% of revenue; if signed, the first Chinese frontier model enters your current cloud contract without you changing a thing.

Watch the 1-minute video