The Morning Build for August 26, 2026: Jalapeño and Groq 3 LPX Benchmark Claims, Perplexity Goes Fully Local, Gemini for Legal Preview
Hardware and deployment moves dominate today: OpenAI and Nvidia showed new inference accelerators with benchmark claims, Perplexity launched a fully local agent appliance with Nvidia hardware, and Google published a preview of Gemini for legal work. These stories focus on inference efficiency, local execution, and industry-specific connectors.
SemiAnalysis lab benchmarks show OpenAI’s Jalapeño leads perf/W vs Blackwell and Rubin A0
- What happened: SemiAnalysis published in-lab InferenceX runs supplied by OpenAI showing Jalapeño A0 engineering samples outperform Nvidia Blackwell and earlier GB200 results on tokens per watt and tokens per user for several open models; OpenAI provided the systems and verified runs in person. SemiAnalysis noted no AgentX multiturn/long-context runs were available and that numbers derive from single-token prediction runs.
- Why it matters: Engineers get a vendor-verified datapoint that a first-generation, inference-focused ASIC can exceed current GPU perf/W on 8k1k-style workloads, highlighting power-constrained datacenter economics and the impact of hardware-software codesign on token throughput per megawatt.
- Outlook: AgentX-style multiturn and long-context benchmark runs are the next public verification step named by SemiAnalysis to assess real-world agentic serving behavior.
Sources: newsletter.semianalysis.com · the-decoder.com
OpenAI says Jalapeño will deploy in small volumes by end of 2026 with larger rollouts in 2027
- What happened: OpenAI showed detailed Jalapeño architecture and benchmark claims at Hot Chips and told press the chip was developed with Broadcom, optimized for minimizing data movement and prefill/communication delays, and estimated small-volume deployment at the end of 2026 with broader deployment in 2027.
- Why it matters: Teams planning inference capacity must factor in a potential new supplier and a multigenerational OpenAI platform that targets perf/W and local KV cache placement, which could alter cost and latency trade-offs once deployed.
- Outlook: End of 2026 small-volume deployments transitioning to larger 2027 rollouts is the timeline OpenAI provided for Jalapeño availability.
Sources: techcrunch.com
Perplexity and Nvidia ship Portable Computer, a fully local agent appliance running on DGX Spark and RTX machines
- What happened: Perplexity launched Portable Computer for Linux today, packaging its Perplexity Computer agent harness, local models (Qwen 3.8 27B and PPLX 27B at launch), inference stack, connectors, and sandboxing to run agents entirely on-device on DGX Spark and RTX GPUs with at least 24GB VRAM; Windows support is scheduled for September and Nemotron Lightning support is coming soon.
- Why it matters: Developers and operators get an appliance-style local execution path where token costs are zero on-device, the harness enforces compact prompts and sandboxing, and hybrid escalation to frontier cloud models is explicitly instrumented, changing cost and data-movement trade-offs for agent workloads.
- Outlook: Windows support arriving in September is the named follow-on milestone for broader Desktop adoption.
Sources: venturebeat.com
Google launches Gemini for Legal preview with MCP connectors to iManage, NetDocuments, DocuSign, Everlaw and Thomson Reuters HighQ
- What happened: Google Cloud announced Gemini Enterprise for Legal as a preview that links Gemini models to legal systems and specialized services via MCP connectors, enables contract review, legal research, and regulatory tracking while respecting access permissions, and is positioned alongside a same-day financial services package and planned healthcare and life sciences offerings.
- Why it matters: Legal and enterprise integrations are now delivered as a previewed, connector-first product that packages pre-built agent ‘skills’ and third-party connectors, which shortens the integration path for firms that need connectored model access to enterprise document systems.
- Outlook: Healthcare and life sciences versions are stated to be on the way as the next product verticals following the legal and financial services previews.
Sources: the-decoder.com
Nvidia moves Groq 3 LPX into production and reports 3,400 tok/s on Gemma 4 31B, with caveats on chip counts
- What happened: Nvidia announced full production for its Groq 3 LPX inference accelerator and published an Artificial Analysis benchmark showing a rack hitting 3,400 tokens per second on Gemma 4 31B over long contexts; coverage notes the result depends on a multi-chip SRAM-heavy dataflow setup requiring at least 64 LPUs for the reported configuration while Cerebras uses fewer accelerators for the same model.
- Why it matters: The LPX design prioritizes high decode throughput via SRAM-local LPUs and rack-level composition, which offers very high token-generation rates for dense models that fit per-rack, but implies different scaling and memory-partition trade-offs compared with large-memory GPU or wafer-scale alternatives.
- Outlook: Nvidia said the LPX will go live later this year, which is the delivery window to watch for availability in cloud and provider offerings.
Sources: the-decoder.com