# What is true on Inference?

> Markdown twin of https://web-production-180a9.up.railway.app/topics/inference · https://web-production-180a9.up.railway.app/topics/inference.md

## On this layer

- org — [NVIDIA](https://web-production-180a9.up.railway.app/companies/nvidia) (12)
- org — OpenAI (6)
- org — [Skild AI](https://web-production-180a9.up.railway.app/companies/skild) (4)
- org — agentic video understanding (3)
- org — [Google](https://web-production-180a9.up.railway.app/companies/google) (3)
- org — tokens processed (3)
- org — AI chip inference cluster (2)
- org — [AMD](https://web-production-180a9.up.railway.app/companies/amd) (2)
- org — Anthropic (2)
- org — MiniMax (2)
- org — Samsung (2)
- org — Vercel (2)
- org — Zhipu (2)
- org — a16z (1)
- org — a16z Machine Age Fund. (1)
- org — ADLINK (1)
- org — Agibot (1)
- org — Apple (1)
- org — Astribot (1)
- org — Axis Robotics (1)
- org — BABA (1)
- org — BABA funding for AI. (1)
- org — Booster Robotics (1)
- org — Broadcom (1)
- org — chemistry (1)
- org — Cloudflare (1)
- org — Cognex (1)
- org — Compute requirements for humanoids. (1)
- org — Current hours of available data for training robots. (1)
- org — Current throughput of valid trajectories collected per hour. (1)
- org — d-Matrix (1)
- org — Deep Robotics (1)
- org — DeepgramAI (1)
- org — development timeline of Moby. (1)
- org — development timeline reduction. (1)
- org — DiscoLoopAI (1)
- org — domestic AI chip inference cluster (1)
- org — DoorDash (1)
- org — Doosan Bobcat (1)
- org — Eptolink (1)
- org — fal (1)
- regulator — [Fcc](https://web-production-180a9.up.railway.app/companies/fcc) (1)
- org — Fetch Robotics (1)
- org — fiftyyears (1)
- org — Flash (1)
- org — Geely Auto (1)
- org — GEN-1.5 can learn new tasks in a few seconds. (1)
- org — General Catalyst (1)
- org — Groq (1)
- org — H3 Max Turbo (1)
- org — HBM prices. (1)
- org — [Hugging Face](https://web-production-180a9.up.railway.app/companies/hugging-face) (1)
- org — inference cost reduction (1)
- org — Innolight (1)
- org — LG Innotek (1)
- org — LimX Dynamics (1)
- org — LM Studio (1)
- org — Lotus Cars (1)
- org — Manycore Tech (1)
- org — MarathonMP (1)
- org — market size for Autonomous AI (1)
- org — Matic (1)
- org — [Meta](https://web-production-180a9.up.railway.app/companies/meta) (1)
- org — Moby deployment in a semiconductor facility. (1)
- org — MRVL (1)
- org — Nebius (1)
- org — Noble Machines (1)
- org — NotionHQ (1)
- org — Nscale (1)
- org — Nscale IPO. (1)
- org — Number of global contributors and collected trajectories on the platform. (1)
- org — Ollama (1)
- org — outsetcap (1)
- org — performance metrics. (1)
- org — Perplexity (1)
- org — power consumption comparison of Jetson Orin Nano 2 (1)
- org — Rockchip (1)
- org — Runway (1)
- org — Salesforce (1)
- org — SambaNova (1)
- org — Samsung Electro Mechanics (1)
- org — Schaeffler (1)
- regulator — sec (1)
- org — Seeed (1)
- org — Simulated robot bodies (1)
- org — SK Hynix (1)
- org — SKHY Indiana packaging plant. (1)
- org — Solomon (1)
- org — success rate of the post-trained policy. (1)
- org — Target valid hours of data per month from the Mobile Ego-Centric App. (1)
- org — tasks available in RoboLab. (1)
- org — [Tesla](https://web-production-180a9.up.railway.app/companies/tesla) (1)
- org — Training Data (1)
- org — training needed for current VLA models to match S1's accuracy (1)
- org — TRON 2 (1)
- org — [Unitree](https://web-production-180a9.up.railway.app/companies/unitree) (1)
- org — validated run uses 64 nodes of 4× GB200 for 60K iterations, roughly 68 hours (17.4K GB200-hours). (1)
- org — Views on the post (1)
- org — [vLLM](https://web-production-180a9.up.railway.app/companies/vllm) (1)
- org — wafer (1)
- org — Wafer's funding announcement. (1)
- org — Waymo (1)
- org — Wing (1)
- org — Wing_VC (1)
- org — Yangtze OF (1)
- org — ycombinator (1)
- org — YMTC (1)
- org — ZiNovaLabs (1)
- org — Zite (1)

## Who captures

- NVIDIA (10)
- @smsehy (6)
- @stretchcloud (4)
- OpenAI (4)
- Skild AI (3)
- MiniMax (2)
- Google (2)
- Zhipu (2)
- Meta (1)
- AMD (1)
- Broadcom (1)
- Tesla (1)
- @SemiAnalysis_ (1)
- Astribot (1)
- Vercel (1)
- @rimtoln (1)
- Zite (1)
- @MelvinInvests (1)
- Rockchip (1)
- @OperationsPLS (1)
- General Catalyst (1)
- @ricci_nov (1)
- Perplexity (1)
- Waymo (1)
- ZiNovaLabs (1)
- wafer (1)
- @tphuang (1)
- @frontrunvc (1)
- https://x.com/LerrelPinto (1)
- Axis Robotics (1)
- https://x.com/Majumdar_Ani (1)

## Sourced numbers

- **H3 Max Turbo:** $0.01
  - _"a $0.01/s video generation endpoint makes AI video economically viable for things like product demo generation, marketing content pipelines, and real-time synthetic data."_
  - Source: [@stretchcloud on X](https://x.com/stretchcloud/status/2095304264994177524)
- **Compute requirements for humanoids.:** 0 units
  - _"Sizing on-chip memory buffers to prevent bus saturation during continuous high-rate tactile policy execution directly dictates compute bill of materials."_
  - Source: [@smsehy on X](https://x.com/smsehy/status/2095050686904062063)
- **TRON 2:** 1 units
  - _"TRON 2 serves as a modular and extensible embodied robotic platform for multi-tool, multi-step tasks across large workspaces."_
  - Source: [LimX Dynamics on X: "Together with ZINOVA's Tool Intelligence, TRON 2 takes on increasingly complex construction workflows.

@ZiNovaLabs builds on TRON 2 to explore an innovative robotic configuration for construction, demonstrating key tasks in a scaled-down tilt-up construction workflow, includi… / X](https://x.com/chris_j_paxton/status/2095114381818253768)
- **Moby deployment in a semiconductor facility.:** 3 units
  - _"bring two Moby3 units to a semiconductor facility for material-handling workflows."_
  - Source: [Noble Machines Accelerates Humanoid Robot Development 3X | NVIDIA](https://x.com/NVIDIARobotics/status/2094923298589024698)
- **development timeline of Moby.:** $3
  - _"Accelerate Moby's development timeline by nearly 3x, from an expected four years to 18 months."_
  - Source: [Noble Machines Accelerates Humanoid Robot Development 3X | NVIDIA](https://x.com/NVIDIARobotics/status/2094923298589024698)
- **development timeline reduction.:** 18 hrs
  - _"from an initial estimated timeline of four years with 50 people to 18 months with 15 people."_
  - Source: [Noble Machines Accelerates Humanoid Robot Development 3X | NVIDIA](https://x.com/NVIDIARobotics/status/2094923298589024698)
- **Wafer's funding announcement.:** $40M
  - _"We’re excited to announce that we’ve raised a $40M Series A!"_
  - Source: [wafer on X: "We’re excited to announce that we’ve raised a $40M Series A!

co-led by @MarathonMP and @chemistry, with participation from @Wing_VC, @AMD Ventures, @outsetcap, @fiftyyears, and @ycombinator, and our existing investors doubling down on @wafer_ai. we are also joined by an incredible lis… / X](https://x.com/frontrunvc/status/2094841271244247488)
- **Google DeepMind's RT-2 model:** 2 units
  - _"it achieved nearly double the performance on novel tasks compared to its predecessor, RT-1."_
  - Source: [smartfitguide on X: "Google DeepMind's RT-2 model treats robot actions as text tokens. By co-fine-tuning pre-trained Vision-Language Models with robotic data, it achieved nearly double the performance on novel tasks compared to its predecessor, RT-1. #AI" / X](https://x.com/bareani21645/status/2094613978739593316)
- **AI chip inference cluster:** 100,000 units
  - _"It already has 100k domestic AI chip inference cluster & plan to increase significantly over next 6 months."_
  - Source: [tphuang on X: "GLM-6 disclosed - more parameter &amp; lower activation ratio &amp; inference cost &amp; improved post training expected. Model to make self training decisions

Expecting significant increased domestic compute addition in next 3 to 6 months. Using self owned, leased &amp; purchased c… / X](https://x.com/tphuang/status/2094601178684264502)
- **tokens processed:** 62,000,000 units
  - _"It did 62T tokens in 6 days w/ Ox Alpha"_
  - Source: [tphuang on X: "GLM-6 disclosed - more parameter &amp; lower activation ratio &amp; inference cost &amp; improved post training expected. Model to make self training decisions

Expecting significant increased domestic compute addition in next 3 to 6 months. Using self owned, leased &amp; purchased c… / X](https://x.com/tphuang/status/2094601178684264502)
- **domestic AI chip inference cluster:** 100,000 units
  - _"It already has 100k domestic AI chip inference cluster & plan to increase significantly over next 6 months."_
  - Source: [@tphuang on X](https://x.com/tphuang/status/2094601174422855998)
- **tokens processed:** 62,000 units
  - _"It did 62T tokens in 6 days w/ Ox Alpha"_
  - Source: [@tphuang on X](https://x.com/tphuang/status/2094601174422855998)
- **inference cost reduction:** $80
  - _"inference cost has dropped 80% since start of yr"_
  - Source: [@tphuang on X](https://x.com/tphuang/status/2094601174422855998)
- **AI chip inference cluster:** 100,000 units
  - _"It already has 100k domestic AI chip inference cluster & plan to increase significantly over next 6 months."_
  - Source: [tphuang on X: "How it views its own models in terms of the kind of tasks them can do. GLM-5.3 in terms of long time horizon work is still short of 1 wk, need at least 1 month to be autonomous.

Flash is the 1st one to use domestic chips &amp; have low cost + intelligence to support long tasks." / X](https://x.com/tphuang/status/2094601170102747138)
- **tokens processed:** 62,000 units
  - _"It did 62T tokens in 6 days w/ Ox Alpha"_
  - Source: [tphuang on X: "How it views its own models in terms of the kind of tasks them can do. GLM-5.3 in terms of long time horizon work is still short of 1 wk, need at least 1 month to be autonomous.

Flash is the 1st one to use domestic chips &amp; have low cost + intelligence to support long tasks." / X](https://x.com/tphuang/status/2094601170102747138)
- **market size for Autonomous AI:** $30,000,000,000
  - _"last step having 30T mkt size according to Show more"_
  - Source: [tphuang on X: "How it views its own models in terms of the kind of tasks them can do. GLM-5.3 in terms of long time horizon work is still short of 1 wk, need at least 1 month to be autonomous.

Flash is the 1st one to use domestic chips &amp; have low cost + intelligence to support long tasks." / X](https://x.com/tphuang/status/2094601170102747138)
- **Funding for Skild AI:** $1.7billion
  - _"Skild has raised nearly $1.7 billion since its founding in 2023 to develop a general-purpose robot brain."_
  - Source: [The Robot Report: Skild AI unveils S1 robot foundation model | AI Understanding](https://x.com/aiuorg/status/2094602992565522608)
- **agentic video understanding:** $66
  - _"reduces costs by up to 66%"_
  - Source: [Introducing Agentic Video in Gemini](https://x.com/GoogleDeepMind/status/2094840182457422260)
- **agentic video understanding:** 88 units
  - _"cuts token consumption by up to 88%"_
  - Source: [Introducing Agentic Video in Gemini](https://x.com/GoogleDeepMind/status/2094840182457422260)
- **agentic video understanding:** 7 units
  - _"boosts quality by up to 7%"_
  - Source: [Introducing Agentic Video in Gemini](https://x.com/GoogleDeepMind/status/2094840182457422260)
- **performance metrics.:** 0 units
  - _"Tokens Per Megawatt"_
  - Source: [@SemiAnalysis_ on X](https://x.com/SemiAnalysis_/status/2094198478331416747)
- **BABA funding for AI.:** $10.2B
  - _"$BABA selling $10.2B of stock to fund AI."_
  - Source: [@ParadisLabs on X](https://x.com/ParadisLabs/status/2094086421413814627)
- **a16z Machine Age Fund.:** $1.1B
  - _"a16z launched $1.1B Machine Age Fund for chips, memory, networking, data centres, cooling, robotics and edge."_
  - Source: [@ParadisLabs on X](https://x.com/ParadisLabs/status/2094086421413814627)
- **Nscale IPO.:** 3 units
  - _"Nscale (backed by $NVDA) targeting up to $3B in US IPO soon."_
  - Source: [@ParadisLabs on X](https://x.com/ParadisLabs/status/2094086421413814627)
- **HBM prices.:** 50 hrs
  - _"HBM contract prices maybe +50% or more in 2027 on the back of the $NVDA server hike = memory makers keep printing earnings."_
  - Source: [@ParadisLabs on X](https://x.com/ParadisLabs/status/2094086421413814627)
- **SKHY Indiana packaging plant.:** 458 units
  - _"$458M CHIPS funding done."_
  - Source: [@ParadisLabs on X](https://x.com/ParadisLabs/status/2094086421413814627)
- **Views on the post:** 267,700 units
  - _"267.7KViews"_
  - Source: [Lerrel Pinto on X: "Turns out that doing In-context learning for robots is not that hard..." / X](https://x.com/LerrelPinto/status/2093400831941116254)
- **Current hours of available data for training robots.:** 2000 hrs
  - _"Today, the industry has roughly 2,000 hours, the best public dataset, Open X-Embodiment."_
  - Source: [Cicada Market Making on X: "https://t.co/26k9SK56mf" / X](https://x.com/axisrobotics/status/2093684446465798629)
- **Number of global contributors and collected trajectories on the platform.:** 150,000 units
  - _"As of late August 2026, the platform has surpassed 150,000 global contributors and collected over 3.7 million trajectories across 4,000+ published tasks."_
  - Source: [Cicada Market Making on X: "https://t.co/26k9SK56mf" / X](https://x.com/axisrobotics/status/2093684446465798629)
- **Current throughput of valid trajectories collected per hour.:** 10,000 units
  - _"Throughput already reaches 10,000 valid trajectories per hour today;"_
  - Source: [Cicada Market Making on X: "https://t.co/26k9SK56mf" / X](https://x.com/axisrobotics/status/2093684446465798629)
- **Target valid hours of data per month from the Mobile Ego-Centric App.:** 10,000 units
  - _"Launching in September 2026 with a target of 10,000+ valid hours of data per month;"_
  - Source: [Cicada Market Making on X: "https://t.co/26k9SK56mf" / X](https://x.com/axisrobotics/status/2093684446465798629)
- **Simulated robot bodies:** 100,000 units
  - _"They trained locomotion across 100,000 simulated robot bodies, which forces the policy to deal with different morphologies instead of just memorizing how one specific robot moves."_
  - Source: [@0xconglomerate on X](https://x.com/0xconglomerate/status/2092554990829379730)
- **training needed for current VLA models to match S1's accuracy:** 100 hrs
  - _"current VLA models would need to be post-trained with 50-100 hours of data collection followed by fine-tuning!"_
  - Source: [Skild AI on X: "Introducing S1, our new foundation model that learns from one example.

It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning.

Watch S1 operate in real-time via in-context learning:" / X](https://x.com/SkildAI/status/2092300842900865389)
- **power consumption comparison of Jetson Orin Nano 2:** $40
  - _"consumes 40% less power at the same performance."_
  - Source: [NVIDIA Announces Jetson Orin Nano 2 Robotics Computer to Redefine Entry-Level Edge AI | NVIDIA Newsroom](https://x.com/NVIDIARobotics/status/2092267694565249127)
- **GEN-1.5 can learn new tasks in a few seconds.:** 1 units
  - _"Introducing GEN-1.5, a one-shot learner."_
  - Source: [Anirudha Majumdar on X: "Very cool work, Pete! I'm curious if you've tried an extreme version of this: no in-context example at all; just place objects in front of the robot and see if it can infer the task. 

I suspect this will have non-trivial success rates, which would also allow you to figure … / X](https://x.com/peteflorence/status/2090592251286335962)
- **validated run uses 64 nodes of 4× GB200 for 60K iterations, roughly 68 hours (17.4K GB200-hours).:** 68 hrs
  - _"Plan compute accordingly."_
  - Source: [Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control | NVIDIA Technical Blog](https://x.com/NVIDIARobotics/status/2095179993974260168)
- **tasks available in RoboLab.:** 120 units
  - _"It executes each action chunk in physics and streams rendered observations back for a true closed loop across 120 language-conditioned manipulation tasks."_
  - Source: [Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control | NVIDIA Technical Blog](https://x.com/NVIDIARobotics/status/2095179993974260168)
- **success rate of the post-trained policy.:** 22.9 units
  - _"In closed-loop RoboLab evaluation, the post-trained Edge policy reaches 22.9% success across 120 language-conditioned manipulation tasks."_
  - Source: [Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control | NVIDIA Technical Blog](https://x.com/NVIDIARobotics/status/2095179993974260168)
- **Training Data:** 20000 hrs
  - _"N1.7 is pretrained on 20K hours of EgoScale human video data alongside diverse robot demonstrations."_
  - Source: [NVIDIA/Isaac-GR00T](https://github.com/NVIDIA/Isaac-GR00T)

## Pulse

- [@stretchcloud on X](https://x.com/stretchcloud/status/2095739885772554382)
  Reef is one of the most ambitious open-source AI infrastructure projects I have seen this year. ~300 GitHub stars in two days. The core idea is genuinely different from anything else in the space: mos…
- [@stretchcloud on X](https://x.com/stretchcloud/status/2095734852259881296)
  What NVIDIA shipped at IFA yesterday quietly changes local inference infrastructure. PAIR (Personal AI Router) is a free, open-source tool that auto-discovers compatible GPUs on your local network and…
- [@tchsignal on X](https://x.com/tchsignal/status/2095733547231514841)
  Meta Puts a Price on AI Interaction Data Meta is offering a striking trade-off for developers using its Muse Spark AI model: much cheaper API access if they opt in to Contributor terms that allow thei…
- [@SemiAnalysis_ on X](https://x.com/SemiAnalysis_/status/2095731112207061029)
  On total tokens per $ TCO, a new AMD MI355x submission beats B300 at lower interactivity ranges on AgentX. Both high throughput configs use disaggregated setups. Shoutout to vLLM, AMD, and LMCache eng…
- [@ricci_nov on X](https://x.com/ricci_nov/status/2095727475540271308)
  Broadcom's core switching GM Asad Khamisy told Semicon Taiwan that co-packaged optics is not necessarily the best answer for scale-up networking, and fits scale-out better. The optics story splits by …
- [@elonmusk on X](https://x.com/elonmusk/status/2095717295192486239)
  Cybercab safety https://t.co/9DQYsSTGeU
- [@smsehy on X](https://x.com/smsehy/status/2095711912843694167)
  A 15-gigawatt deficit in grid power will accelerate the shift toward embedded edge intelligence and model quantization. Running efficient inference directly inside vehicle gateways and factory edge no…
- [@smsehy on X](https://x.com/smsehy/status/2095711321790566803)
  Physical constraints across high-bandwidth memory packaging and behind-the-meter power interconnects are forcing compute optimization down to the edge. When centralized megawatts take four years to co…
- [@stretchcloud on X](https://x.com/stretchcloud/status/2095705156671480096)
  The 30-50x cost difference between computer use and MCP keeps rattling around in my head. The framing that clicks for me: computer use is paying for a human-in-the-loop at model prices. Every screen i…
- [@stretchcloud on X](https://x.com/stretchcloud/status/2095700123464733012)
  The thing I keep noticing across agentic frameworks in 2026: resumability is now the reliability primitive everyone is building toward. Genkit Go 1.13 ships it properly. A Generate call that fails at …
- [@DrJimFan on X](https://x.com/DrJimFan/status/2095696882844463382)
  Good old days at OpenAI in 2016: an agent stares at screen pixels, moves a mouse, and books a flight on United. We called it World of Bits, inside OpenAI Universe. 10 yrs later, Astra is reincarnated …
- [@rimtoln on X](https://x.com/rimtoln/status/2095628517103071435)
  GPT-6 ASTRA ISN'T A PATCH. IT'S A GENERATION FLIP. openai just shipped the model they call a generational leap past gpt-5.6 sol not better chat better computer use · coding · cyber · science ▹ why thi…
- [@SemiAnalysis_ on X](https://x.com/SemiAnalysis_/status/2095595233064972516)
  Shoutout to the cracked team at @vllm_project that implemented recent agentic workload optimizations. (1/5)🧵 https://t.co/BsXCtLBwvU
- [@XRoboHub on X](https://x.com/XRoboHub/status/2095588510950727820)
  Robots can’t stop while the model thinks. In a high-speed throw, even one pause can kill the momentum. Astribot released SmoothRL for online RL during async inference. S1 keeps moving as the model com…
- [@stretchcloud on X](https://x.com/stretchcloud/status/2095578824046219594)
  One command now routes Claude Code, Cursor, Codex, Cline, and four other coding harnesses through a single AI Gateway. That is what Vercel's setup flow does: run `vercel ai-gateway coding-agents setup…
- [@rimtoln on X](https://x.com/rimtoln/status/2095568388684738657)
  EVERYONE SHOWS THE 300-AGENT SWARM. almost nobody shows what happens after 300 agents are easy to spawn turning 300 outputs into one structure is the hard part ▹ the pipeline nobody demos collect → co…
- [@stretchcloud on X](https://x.com/stretchcloud/status/2095565486230782032)
  The gap between "AI can help you build this" and "AI can actually deploy this for you" just got smaller. Zite released an MCP server that lets Claude, ChatGPT, and Cursor generate fully deployable bus…
- [@tchsignal on X](https://x.com/tchsignal/status/2095561122074198141)
  Nvidia RTX Spark N1X Brings Up to 128GB of Unified Memory to Windows PCs Nvidia's RTX Spark N1X is more than another increase in CPU and GPU core counts. The new platform combines a Grace CPU, Blackwe…
- [@tchsignal on X](https://x.com/tchsignal/status/2095558659279605842)
  Nvidia PAIR Turns Multiple Computers Into a Local AI Compute Pool Nvidia has launched the beta of Personal AI Router, or PAIR, a free software tool designed to make multiple compatible computers on th…
- [@tchsignal on X](https://x.com/tchsignal/status/2095551413661335854)
  Nvidia’s $12.9B Hugging Face Deal Raises a Bigger Question About Open AI Nvidia has agreed to acquire Hugging Face in a deal valued at about $12.93 billion, but the transaction has not closed yet. It …
- [@MelvinInvests on X](https://x.com/MelvinInvests/status/2095548651619877295)
  AI’s biggest bottleneck is moving data and that could still create huge opportunities for optical networking companies (Save this) The chart shows a 1.6T optical transceiver, a device that transfers d…
- [@stretchcloud on X](https://x.com/stretchcloud/status/2095494015567458323)
  Video generation just crossed a threshold I didn't expect until 2027. https://t.co/wgZmOqI59X took MiniMax's H3 model, post-trained it for quality and cost, then ran it through a custom inference engi…
- [@smsehy on X](https://x.com/smsehy/status/2095493512171028642)
  An autonomous vehicle generates 1 to 2 TB of sensor data per hour. Streaming that to the cloud for inference costs more than the compute hardware in the vehicle. The economics of edge AI in automotive…
- [@smsehy on X](https://x.com/smsehy/status/2095441064119373989)
  Enterprise customers deploying autonomous physical fleets cannot rely on centralized multi-tenant cloud APIs. Full operational sovereignty requires edge inference execution completely decoupled from p…
- [@smsehy on X](https://x.com/smsehy/status/2095435750141829216)
  In automotive zonal compute, reserving leading-edge nodes for central AI inference while keeping I/O and power interfaces on mature 28nm silicon is the only viable path to maintain AEC-Q100 thermal re…
- [@seeedstudio on X](https://x.com/seeedstudio/status/2095433547595325555)
  🔥🔥NEW Launch: reComputer RK3576 Dev Kit/Module/IO Board is Officially HERE!!!! Level up your edge AI development with Rockchip #RK3576 today! 🎉 🧩🧩Dev Kit / Module / IO Board: https://t.co/xQVIT62…
- [@OperationsPLS on X](https://x.com/OperationsPLS/status/2095425090670449005)
  Edge AI + Physical AI is the architecture shift. Moving inference to the robot eliminates cloud latency — critical for real-time manipulation and safety. HBM5 at the edge means robots think at the poi…
- [@stretchcloud on X](https://x.com/stretchcloud/status/2095334463802851765)
  Sodium v2 introduces a WebMCP layer. Any website can now serve structured data directly to an AI agent without HTML scraping. This is AEO (Agent Engine Optimization) in practice. The web is starting t…
- [@stretchcloud on X](https://x.com/stretchcloud/status/2095309298138272015)
  12 PhDs and a Fields Medal winner built this in 4 months, backed by General Catalyst. https://t.co/pyhyV4u6I4 has a protocol that passes hidden states directly from a 753B frontier model to a 4B edge …
- [@stretchcloud on X](https://x.com/stretchcloud/status/2095304264994177524)
  The pattern in AI inference is consistent: the fast-and-cheap version of a frontier model ships 2-4 weeks after the frontier version. H3 Max Turbo confirms this for video generation.H3 Max launched on…
- [@ricci_nov on X](https://x.com/ricci_nov/status/2095304089353834603)
  HBQ: W4A5 matching W4A16 accuracy, in less silicon area than NVFP4. On a 28nm ASIC it shows 2.3x area and 4.6x energy efficiency over weight-only quantization, and 1.6-3.3x lower system energy than pr…
- [@stretchcloud on X](https://x.com/stretchcloud/status/2095294198886903921)
  The hybrid compute pattern is getting serious. Perplexity just open-sourced Lily, the local inference engine powering their Mac hybrid compute feature, and the architecture choices are worth understan…
- [@smsehy on X](https://x.com/smsehy/status/2095293945534091544)
  Waymo designed its own inference ASIC rather than using off-the-shelf GPUs. The thermal envelope inside a vehicle roof module does not allow a 300W GPU. Automotive compute is a packaging and thermal p…
- [@smsehy on X](https://x.com/smsehy/status/2095050686904062063)
  Localized low-latency inference on humanoids requires automotive-grade high-bandwidth memory that withstands intense mechanical vibration and thermal cycling. Sizing on-chip memory buffers to prevent …
- [LimX Dynamics on X: "Together with ZINOVA's Tool Intelligence, TRON 2 takes on increasingly complex construction workflows.

@ZiNovaLabs builds on TRON 2 to explore an innovative robotic configuration for construction, demonstrating key tasks in a scaled-down tilt-up construction workflow, includi… / X](https://x.com/chris_j_paxton/status/2095114381818253768)
  Post Log in Sign up Post LimX Dynamics on X: "Together with ZINOVA's Tool Intelligence, TRON 2 takes on increasingly complex construction workflows. @ZiNovaLabs builds on TRON 2 to explore an innovati…
- [Noble Machines Accelerates Humanoid Robot Development 3X | NVIDIA](https://x.com/NVIDIARobotics/status/2094923298589024698)
  Manufacturing \| Robotics Noble Machines Accelerates General Purpose Industrial Robot Development With NVIDIA Isaac GR00T Learn More Objective Noble Machines is advancing industrial robot development …
- [Seeed reBot Arm B601-RS: Physical AI & VLA Model Course | NVIDIA DLI Series](https://x.com/NVIDIARobotics/status/2094850247184826558)
  Warehouse China Warehouse US Warehouse Germany Warehouse The store will not work correctly in the case when cookies are disabled. US Warehouse: Enjoy FREE UNIUNI shipping on orders under $50! (Exclude…
- [wafer on X: "We’re excited to announce that we’ve raised a $40M Series A!

co-led by @MarathonMP and @chemistry, with participation from @Wing_VC, @AMD Ventures, @outsetcap, @fiftyyears, and @ycombinator, and our existing investors doubling down on @wafer_ai. we are also joined by an incredible lis… / X](https://x.com/frontrunvc/status/2094841271244247488)
  Post Log in Sign up Post wafer on X: "We’re excited to announce that we’ve raised a $40M Series A! co-led by @MarathonMP and @chemistry, with participation from @Wing\VC, @AMD Ventures, @outsetcap, @f…
- [smartfitguide on X: "Google DeepMind's RT-2 model treats robot actions as text tokens. By co-fine-tuning pre-trained Vision-Language Models with robotic data, it achieved nearly double the performance on novel tasks compared to its predecessor, RT-1. #AI" / X](https://x.com/bareani21645/status/2094613978739593316)
  Post Log in Sign up Post smartfitguide on X: "Google DeepMind's RT-2 model treats robot actions as text tokens. By co-fine-tuning pre-trained Vision-Language Models with robotic data, it achieved near…
- [tphuang on X: "GLM-6 disclosed - more parameter &amp; lower activation ratio &amp; inference cost &amp; improved post training expected. Model to make self training decisions

Expecting significant increased domestic compute addition in next 3 to 6 months. Using self owned, leased &amp; purchased c… / X](https://x.com/tphuang/status/2094601178684264502)
  Post Log in Sign up Post tphuang on X: "GLM-6 disclosed - more parameter &amp; lower activation ratio &amp; inference cost &amp; improved post training expected. Model to make self training decisions …
- [@tphuang on X](https://x.com/tphuang/status/2094601174422855998)
  It already went oversea w/ Malaysia deployment & will have more oversea CSP cooperation deal signed in next month or 2. It already has 100k domestic AI chip inference cluster & plan to increase signif…
- [tphuang on X: "How it views its own models in terms of the kind of tasks them can do. GLM-5.3 in terms of long time horizon work is still short of 1 wk, need at least 1 month to be autonomous.

Flash is the 1st one to use domestic chips &amp; have low cost + intelligence to support long tasks." / X](https://x.com/tphuang/status/2094601170102747138)
  Post Log in Sign up Post tphuang on X: "How it views its own models in terms of the kind of tasks them can do. GLM-5.3 in terms of long time horizon work is still short of 1 wk, need at least 1 month …
- [The Robot Report: Skild AI unveils S1 robot foundation model | AI Understanding](https://x.com/aiuorg/status/2094602992565522608)
  Back to News ProductAI Understanding briefing The Robot Report: Skild AI unveils S1 robot foundation model The Robot Report reports that Skild AI unveiled S1, a robot foundation model that the company…
- [Introducing Agentic Video in Gemini](https://x.com/GoogleDeepMind/status/2094840182457422260)
  Introducing agentic video understanding with Gemini Sep 01, 2026 \| 7 min read - x.com - Facebook - LinkedIn - Mail - Copy link Our new agentic feature for video analysis cuts token consumption by up …
- [@SemiAnalysis_ on X](https://x.com/SemiAnalysis_/status/2094198478331416747)
  Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) This week Bryan, Myron and Jordan (@JordanNanos) discuss our recent article on OpenAI Jalapeño. They cover the performance, archi…
- [@ParadisLabs on X](https://x.com/ParadisLabs/status/2094086421413814627)
  Some AI/semis notes from the past week: -> $NVDA huge Q2 revenue and profit growth. 2027+ is more growth even under a bottlenecked environment. Nvidia GOAT company? -> $NBIS first cloud customer for $…
- [Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) - YouTube](https://x.com/SemiAnalysis_/status/2093715936406610233)
  Error 401 (Bad Request)!!1 401. That’s an error. The server cannot process the request because it is malformed. It should not be retried. That’s all we know. Back Skip navigation Search Search with yo…
- [@frontrunvc on X](https://x.com/frontrunvc/status/2093420293188391334)
  plugged the Machine Age memo into claude + the @frontrunvc mcp. companies flagged in the last 30 days that fit the thesis: 👇 chips + compute: @stanmachines AI that designs chips, YC S26 @neurophos a …
- [Lerrel Pinto on X: "Turns out that doing In-context learning for robots is not that hard..." / X](https://x.com/LerrelPinto/status/2093400831941116254)
  Post Log in Sign up Post Lerrel Pinto on X: "Turns out that doing In-context learning for robots is not that hard..." - Lerrel Pinto @LerrelPinto Turns out that doing In-context learning for robots is…
- [Cicada Market Making on X: "https://t.co/26k9SK56mf" / X](https://x.com/axisrobotics/status/2093684446465798629)
  Post Log in Sign up Post - Cicada Market Making @cicada\mm Inside Robotics and Physical AI Featuring insights from Axis Robotics In 2026, the artificial intelligence industry learned to solve the comp…
- [@0xconglomerate on X](https://x.com/0xconglomerate/status/2092554990829379730)
  Is it really possible for a robot to learn a new task from one human video, with zero fine-tuning? I’ve read a lot of hype around Skild’s S1, with some already calling it robotics’ “ChatGPT moment.” T…
- [Skild AI on X: "Introducing S1, our new foundation model that learns from one example.

It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning.

Watch S1 operate in real-time via in-context learning:" / X](https://x.com/SkildAI/status/2092300842900865389)
  Post Log in Sign up Post Skild AI on X: "Introducing S1, our new foundation model that learns from one example. It can be taught 10-minute long tasks that it has never seen before, from one video prom…
- [NVIDIA Announces Jetson Orin Nano 2 Robotics Computer to Redefine Entry-Level Edge AI | NVIDIA Newsroom](https://x.com/NVIDIARobotics/status/2092267694565249127)
  NVIDIA Announces Jetson Orin Nano 2 Robotics Computer to Redefine Entry-Level Edge AI August 25, 2026 News Summary: - NVIDIA Jetson Orin Nano 2, a new robotics computer for entry-level edge AI, enable…
- [Anirudha Majumdar on X: "Very cool work, Pete! I'm curious if you've tried an extreme version of this: no in-context example at all; just place objects in front of the robot and see if it can infer the task. 

I suspect this will have non-trivial success rates, which would also allow you to figure … / X](https://x.com/peteflorence/status/2090592251286335962)
  Post Log in Sign up Post Anirudha Majumdar on X: "Very cool work, Pete! I'm curious if you've tried an extreme version of this: no in-context example at all; just place objects in front of the robot a…
- [Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control | NVIDIA Technical Blog](https://x.com/NVIDIARobotics/status/2095179993974260168)
  Technical Blog Subscribe Related Resources Robotics Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control Aug 19, 2026 By Saeed Babamohamadi +12 Like Discuss (0) - L - T - F - R - E AI-Generated…
- [NVIDIA/Isaac-GR00T](https://github.com/NVIDIA/Isaac-GR00T)
  <div align="center"> <img src="media/header_compress.png" width="800" alt="NVIDIA Isaac GR00T N1.7 Header"> <!-- --- --> <p style="font-size: 1.2em;"> <a href="https://developer.nvidia.com/isaac/gr00t…

---

Source: Robotics insights · https://web-production-180a9.up.railway.app/topics/inference.md