NVIDIA12OpenAI6Skild AI4agentic video understanding3Google3tokens processed3AI chip inference cluster2AMD2Anthropic2MiniMax2Samsung2Vercel2Zhipu2a16z1a16z Machine Age Fund.1ADLINK1Agibot1Apple1Astribot1Axis Robotics1BABA1BABA funding for AI.1Booster Robotics1Broadcom1chemistry1Cloudflare1Cognex1Compute requirements for humanoids.1Current hours of available data for training robots.1Current throughput of valid trajectories collected per hour.1d-Matrix1Deep Robotics1DeepgramAI1development timeline of Moby.1development timeline reduction.1DiscoLoopAI1domestic AI chip inference cluster1DoorDash1Doosan Bobcat1Eptolink1fal1Fcc1Fetch Robotics1fiftyyears1Flash1Geely Auto1GEN-1.5 can learn new tasks in a few seconds.1General Catalyst1Groq1H3 Max Turbo1HBM prices.1Hugging Face1inference cost reduction1Innolight1LG Innotek1LimX Dynamics1LM Studio1Lotus Cars1Manycore Tech1MarathonMP1market size for Autonomous AI1Matic1Meta1Moby deployment in a semiconductor facility.1MRVL1Nebius1Noble Machines1NotionHQ1Nscale1Nscale IPO.1Number of global contributors and collected trajectories on the platform.1Ollama1outsetcap1performance metrics.1Perplexity1power consumption comparison of Jetson Orin Nano 21Rockchip1Runway1Salesforce1SambaNova1Samsung Electro Mechanics1Schaeffler1sec1Seeed1Simulated robot bodies1SK Hynix1SKHY Indiana packaging plant.1Solomon1success rate of the post-trained policy.1Target valid hours of data per month from the Mobile Ego-Centric App.1tasks available in RoboLab.1Tesla1Training Data1training needed for current VLA models to match S1's accuracy1TRON 21Unitree1validated run uses 64 nodes of 4× GB200 for 60K iterations, roughly 68 hours (17.4K GB200-hours).1Views on the post1vLLM1wafer1Wafer's funding announcement.1Waymo1Wing1Wing_VC1Yangtze OF1ycombinator1YMTC1ZiNovaLabs1Zite1
H3 Max Turbo$0.01“a $0.01/s video generation endpoint makes AI video economically viable for things like product demo generation, marketing content pipelines, and real-time synthetic data.”@stretchcloud on X
Compute requirements for humanoids.0 units“Sizing on-chip memory buffers to prevent bus saturation during continuous high-rate tactile policy execution directly dictates compute bill of materials.”@smsehy on X
domestic AI chip inference cluster100,000 units“It already has 100k domestic AI chip inference cluster & plan to increase significantly over next 6 months.”@tphuang on X
tokens processed62,000 units“It did 62T tokens in 6 days w/ Ox Alpha”@tphuang on X
inference cost reduction$80“inference cost has dropped 80% since start of yr”@tphuang on X
BABA funding for AI.$10.2B“$BABA selling $10.2B of stock to fund AI.”@ParadisLabs on X
a16z Machine Age Fund.$1.1B“a16z launched $1.1B Machine Age Fund for chips, memory, networking, data centres, cooling, robotics and edge.”@ParadisLabs on X
Nscale IPO.3 units“Nscale (backed by $NVDA) targeting up to $3B in US IPO soon.”@ParadisLabs on X
HBM prices.50 hrs“HBM contract prices maybe +50% or more in 2027 on the back of the $NVDA server hike = memory makers keep printing earnings.”@ParadisLabs on X
SKHY Indiana packaging plant.458 units“$458M CHIPS funding done.”@ParadisLabs on X
Number of global contributors and collected trajectories on the platform.150,000 units“As of late August 2026, the platform has surpassed 150,000 global contributors and collected over 3.7 million trajectories across 4,000+ published tasks.”Cicada Market Making on X: "https://t.co/26k9SK56mf" / X
Simulated robot bodies100,000 units“They trained locomotion across 100,000 simulated robot bodies, which forces the policy to deal with different morphologies instead of just memorizing how one specific robot moves.”@0xconglomerate on X
Reef is one of the most ambitious open-source AI infrastructure projects I have seen this year. ~300 GitHub stars in two days. The core idea is genuinely different from anything else in the space: most RL post-training pipelines treat the model and the scaffolding around it as separate problems. You train one, then you write the other. Reef rejects that separation entirely. It co-evolves model weights and agent harness simultaneously, using live task outcomes as the feedback signal for both. Model side: SAO (Self-Aligned Optimization) adjusts weights from what the agent actually accomplished.…
12:45 AM ET
NVIDIA
$NVDA
12:45 AM ETNVIDIA
@stretchcloud on X
What NVIDIA shipped at IFA yesterday quietly changes local inference infrastructure. PAIR (Personal AI Router) is a free, open-source tool that auto-discovers compatible GPUs on your local network and routes inference requests to whichever machine has capacity. RTX 20-series and above, Apple M4, DGX Spark all supported. Integrates natively with Ollama and LM Studio. A few things to unpack. This is load balancing, not memory pooling. PAIR does not shard a 70B model across three machines or aggregate VRAM. Each machine runs models that fit in its own memory. What PAIR does: intelligently…
12:39 AM ET
Meta
12:39 AM ETMeta
@tchsignal on X
Meta Puts a Price on AI Interaction Data Meta is offering a striking trade-off for developers using its Muse Spark AI model: much cheaper API access if they opt in to Contributor terms that allow their prompts and model outputs to contribute to future model development. Standard pricing is $1.25 per million input tokens and $4.25 per million output tokens. Under Contributor pricing, those rates fall to $0.10 and $0.20 respectively. That distinction matters. Meta is not paying users cash, and the Contributor terms are not the default for every Muse Spark customer. Standard pricing remains…
12:30 AM ET
AMD
12:30 AM ETAMD
@SemiAnalysis_ on X
On total tokens per $ TCO, a new AMD MI355x submission beats B300 at lower interactivity ranges on AgentX. Both high throughput configs use disaggregated setups. Shoutout to vLLM, AMD, and LMCache engineers. https://t.co/Sc35IXkXY8
12:15 AM ET
BRBroadcom
12:15 AM ETBRBroadcom
@ricci_nov on X
Broadcom's core switching GM Asad Khamisy told Semicon Taiwan that co-packaged optics is not necessarily the best answer for scale-up networking, and fits scale-out better. The optics story splits by tier rather than replacing copper everywhere at once. https://t.co/5rPTq1LFXO
11:35 PM ET
Tesla
$TSLA
11:35 PM ETTesla
@elonmusk on X
Cybercab safety https://t.co/9DQYsSTGeU
11:13 PM ET
@smsehy
11:13 PM ET@smsehy
@smsehy on X
A 15-gigawatt deficit in grid power will accelerate the shift toward embedded edge intelligence and model quantization. Running efficient inference directly inside vehicle gateways and factory edge nodes bypasses utility interconnect queues. https://t.co/m8Mi9dmI5E
11:11 PM ET
@smsehy
11:11 PM ET@smsehy
@smsehy on X
Physical constraints across high-bandwidth memory packaging and behind-the-meter power interconnects are forcing compute optimization down to the edge. When centralized megawatts take four years to commission, running quantized models on ruggedized plant controllers is the only operational option.
10:47 PM ET
@stretchcloud
10:47 PM ET@stretchcloud
@stretchcloud on X
The 30-50x cost difference between computer use and MCP keeps rattling around in my head. The framing that clicks for me: computer use is paying for a human-in-the-loop at model prices. Every screen interaction generates a screenshot, feeds it through vision, waits for coordinate output, executes the click, takes another screenshot. You are spending tokens on what amounts to OCR and cursor navigation, round after round. MCP replaces that entire loop with a typed function call. No screenshots. No vision pass. No coordinate uncertainty. The model calls a tool, the tool returns structured data,…
10:27 PM ET
@stretchcloud
10:27 PM ET@stretchcloud
@stretchcloud on X
The thing I keep noticing across agentic frameworks in 2026: resumability is now the reliability primitive everyone is building toward. Genkit Go 1.13 ships it properly. A Generate call that fails at tool round five returns what it finished alongside the error. You pass resp.History() back in, only the failed step reruns. No wasted tool calls. No redone work. The same logic extends to full agent sessions. A failed or cancelled turn saves completed rounds as a snapshot. Send an empty input, it picks up from there. This matters for production. Most AI agent failures today are partial. The agent…
10:14 PM ET
OPOpenAI
10:14 PM ETOPOpenAI
@DrJimFan on X
Good old days at OpenAI in 2016: an agent stares at screen pixels, moves a mouse, and books a flight on United. We called it World of Bits, inside OpenAI Universe. 10 yrs later, Astra is reincarnated in the same universe. Even the naming is astronomically correct 😆 Universe was perhaps the most ambitious AI infra project at the time, but we couldn't quite figure out how to solve it. A policy with zero prior knowledge of what a "submit" button does has to rediscover the entire internet visual lingua by trial and error. In retrospect, RL from scratch against hand-drawn, per-task "artisan"…
5:42 PM ET
OPOpenAI
5:42 PM ETOPOpenAI
@rimtoln on X
GPT-6 ASTRA ISN'T A PATCH. IT'S A GENERATION FLIP. openai just shipped the model they call a generational leap past gpt-5.6 sol not better chat better computer use · coding · cyber · science ▹ why this is the breakthrough first openai model at Critical cyber threshold agent stacks that actually drive the machine browser · forms · repos · multi-step work enterprise-first rollout · daybreak defenders first ▹ the scoreboard (vendor table) frontiermath t4 · astra 97.6 · fable 5.1 87.8 sol was 83.0 · that's a real cliff terminal-bench science · 64.6 vs fable 52.6 automationbench · 41.4 vs 31.4…
3:30 PM ET
@SemiAnalysis_
3:30 PM ET@SemiAnalysis_
@SemiAnalysis_ on X
Shoutout to the cracked team at @vllm_project that implemented recent agentic workload optimizations. (1/5)🧵 https://t.co/BsXCtLBwvU
3:03 PM ET
ASAstribot
3:03 PM ETASAstribot
@XRoboHub on X
Robots can’t stop while the model thinks. In a high-speed throw, even one pause can kill the momentum. Astribot released SmoothRL for online RL during async inference. S1 keeps moving as the model computes the next action chunk. Most actions in a chunk never execute. Train on the full chunk, and RL credits or blames moves that never happened. SmoothRL learns only from executed actions, matching real deployment timing. After 250 rollouts, tossing jumped 39%→94%, pen capping 8%→83%, and box opening 30%→90%. One autonomous toss cut acceleration RMS by 52% and jerk RMS by 47%. The model keeps…
2:25 PM ET
VEVercel
2:25 PM ETVEVercel
@stretchcloud on X
One command now routes Claude Code, Cursor, Codex, Cline, and four other coding harnesses through a single AI Gateway. That is what Vercel's setup flow does: run `vercel ai-gateway coding-agents setup`, it detects which agents you have installed, writes their configuration, and provisions an API key. No markup added on top. 300-plus models from 30-plus providers, behind one endpoint. The practical gain is real. When Anthropic, OpenAI, and SpaceXAI all hit outages in the same window this week, a gateway with automatic fallbacks made the difference between an agent that stalled and one that…
1:43 PM ET
@rimtoln
1:43 PM ET@rimtoln
@rimtoln on X
EVERYONE SHOWS THE 300-AGENT SWARM. almost nobody shows what happens after 300 agents are easy to spawn turning 300 outputs into one structure is the hard part ▹ the pipeline nobody demos collect → connect → organize → unify one context graph at the end the swarm starts as hundreds of isolated research paths then sources start sharing entities links form · contradictions surface weak claims stay exposed eventually the mess collapses into one structure the whole system can reason over ▹ the reframe that's the part you usually never see the swarm is the demo the graph is the product bookmark…
1:32 PM ET
ZIZite
1:32 PM ETZIZite
@stretchcloud on X
The gap between "AI can help you build this" and "AI can actually deploy this for you" just got smaller. Zite released an MCP server that lets Claude, ChatGPT, and Cursor generate fully deployable business applications, not just code files you still have to wire up yourself. The model describes the app, Zite handles the infrastructure and deployment. This is a different category from code generation. Code generation gives you an artifact you own and maintain. Deployment-as-a-tool gives you a running application where the lifecycle management is abstracted away. The distinction matters for the…
1:14 PM ET
Nvidia
$NVDA
1:14 PM ETNvidia
@tchsignal on X
Nvidia RTX Spark N1X Brings Up to 128GB of Unified Memory to Windows PCs Nvidia's RTX Spark N1X is more than another increase in CPU and GPU core counts. The new platform combines a Grace CPU, Blackwell RTX GPU and unified memory in Windows 11 laptops and small desktops. The higher-end configuration pairs 20 CPU cores with 6,144 CUDA cores and supports up to 128GB of LPDDR5X unified memory. A second configuration combines 18 CPU cores with 5,120 CUDA cores and supports up to 64GB according to Nvidia's current specifications. Laptop versions operate within a 45–80W TDP range, while Nvidia…
1:04 PM ET
Nvidia
$NVDA
1:04 PM ETNvidia
@tchsignal on X
Nvidia PAIR Turns Multiple Computers Into a Local AI Compute Pool Nvidia has launched the beta of Personal AI Router, or PAIR, a free software tool designed to make multiple compatible computers on the same local network available for AI inference through a single endpoint. PAIR can route independent requests to available machines based on factors including model availability and GPU utilization. It supports local inference backends including Ollama and LM Studio, with compatible hardware spanning GeForce RTX 20 Series and newer GPUs, RTX PRO systems, DGX Spark and Apple M4 or newer devices.…
12:36 PM ET
Nvidia
$NVDA
12:36 PM ETNvidia
@tchsignal on X
Nvidia’s $12.9B Hugging Face Deal Raises a Bigger Question About Open AI Nvidia has agreed to acquire Hugging Face in a deal valued at about $12.93 billion, but the transaction has not closed yet. It is expected to complete in the first half of 2027, subject to regulatory approvals and other closing conditions. The bigger issue is what Nvidia ownership could mean for an AI platform used by more than 18 million developers, researchers and creators, with over 3 million models, 500,000 datasets and 1 million applications. Nvidia says Hugging Face will remain an open platform. Developers will not…
12:25 PM ET
@MelvinInvests
12:25 PM ET@MelvinInvests
@MelvinInvests on X
AI’s biggest bottleneck is moving data and that could still create huge opportunities for optical networking companies (Save this) The chart shows a 1.6T optical transceiver, a device that transfers data between AI servers, switches, GPUs, and fiber optic networks. 1.6T means it can theoretically move up to 1.6 terabits of data per second, or 1,600 gigabits and that is about twice the speed of an 800G connection. This technology is important because AI data centers contain thousands of GPUs that must constantly exchange information. As AI models become larger, slow connections can leave…
8:48 AM ET
MIMiniMax
8:48 AM ETMIMiniMax
@stretchcloud on X
Video generation just crossed a threshold I didn't expect until 2027. https://t.co/wgZmOqI59X took MiniMax's H3 model, post-trained it for quality and cost, then ran it through a custom inference engine. The result: a 5-second 768p clip renders in about 3 seconds. Faster than real-time. You finish reading the prompt before the video finishes generating. The inference numbers matter because they change the workflow. When generation takes 60+ seconds, you commit to a prompt and wait. When it takes 3 seconds, you iterate. Sketch, reject, sketch again. That shift is what moves video from…
8:46 AM ET
@smsehy
8:46 AM ET@smsehy
@smsehy on X
An autonomous vehicle generates 1 to 2 TB of sensor data per hour. Streaming that to the cloud for inference costs more than the compute hardware in the vehicle. The economics of edge AI in automotive were settled before the debate started.
5:17 AM ET
@smsehy
5:17 AM ET@smsehy
@smsehy on X
Enterprise customers deploying autonomous physical fleets cannot rely on centralized multi-tenant cloud APIs. Full operational sovereignty requires edge inference execution completely decoupled from public network connectivity. https://t.co/YJl2DcWQOa
4:56 AM ET
@smsehy
4:56 AM ET@smsehy
@smsehy on X
In automotive zonal compute, reserving leading-edge nodes for central AI inference while keeping I/O and power interfaces on mature 28nm silicon is the only viable path to maintain AEC-Q100 thermal reliability and control bill of materials cost. https://t.co/a4PfhQqwNR
4:47 AM ET
RORockchip
4:47 AM ETRORockchip
@seeedstudio on X
🔥🔥NEW Launch: reComputer RK3576 Dev Kit/Module/IO Board is Officially HERE!!!! Level up your edge AI development with Rockchip #RK3576 today! 🎉 🧩🧩Dev Kit / Module / IO Board: https://t.co/xQVIT62cbw Powered by #Rockchip #RK3576 octa-core #CPU + 6 TOPS #NPU, this series delivers reliable, low-latency on-device AI inference without cloud dependency, fully empowering industrial edge intelligence, smart vision, local #LLM/#VLM, and offline speech AI applications. 💻🤖 ✅ Powerful Edge AI Performance -YOLO11 at 77.9FPS (640×640) -DeepSeek/Qwen/Whisper -8K video decoding and 4K@60fps encoding…
4:14 AM ET
@OperationsPLS
4:14 AM ET@OperationsPLS
@OperationsPLS on X
Edge AI + Physical AI is the architecture shift. Moving inference to the robot eliminates cloud latency — critical for real-time manipulation and safety. HBM5 at the edge means robots think at the point of action. https://t.co/Z8PWXWTWjM
10:14 PM ET
@stretchcloud
10:14 PM ET@stretchcloud
@stretchcloud on X
Sodium v2 introduces a WebMCP layer. Any website can now serve structured data directly to an AI agent without HTML scraping. This is AEO (Agent Engine Optimization) in practice. The web is starting to build for agents, not just browsers. I built DeepScrape to handle exactly this: giving agents reliable, structured access to web data at inference time. https://t.co/JZYTWY8roq https://t.co/oUD6R0G9j7
8:34 PM ET
GCGeneral Catalyst
8:34 PM ETGCGeneral Catalyst
@stretchcloud on X
12 PhDs and a Fields Medal winner built this in 4 months, backed by General Catalyst. https://t.co/pyhyV4u6I4 has a protocol that passes hidden states directly from a 753B frontier model to a 4B edge model at inference time. No text exchanged between the models. Neither model is fine-tuned. They come from different model families entirely. The result: 80% as accurate as the frontier model alone, running 20x faster. First place on ARC-AGI. The mechanism is what matters here. Most model compression strategies force a choice between accuracy and cost. You either run the big model and accept high…
8:14 PM ET
MIMiniMax
8:14 PM ETMIMiniMax
@stretchcloud on X
The pattern in AI inference is consistent: the fast-and-cheap version of a frontier model ships 2-4 weeks after the frontier version. H3 Max Turbo confirms this for video generation.H3 Max launched on August 26 at $0.08 per second of 768p output. H3 Max Turbo launches today at $0.01 per second, runs at 2x the speed, and hits the 97th percentile of H3 Max quality on evals. Seven days to a 50% cost reduction and 2x throughput improvement.The mechanism matters: fal built a separate inference engine specifically around MiniMax's H3 open-weights model architecture. They're not running H3 Max…
8:13 PM ET
@ricci_nov
8:13 PM ET@ricci_nov
@ricci_nov on X
HBQ: W4A5 matching W4A16 accuracy, in less silicon area than NVFP4. On a 28nm ASIC it shows 2.3x area and 4.6x energy efficiency over weight-only quantization, and 1.6-3.3x lower system energy than prior block quantization. Inference cost is still a datapath problem. https://t.co/08u6ZPO2mO
7:34 PM ET
PEPerplexity
7:34 PM ETPEPerplexity
@stretchcloud on X
The hybrid compute pattern is getting serious. Perplexity just open-sourced Lily, the local inference engine powering their Mac hybrid compute feature, and the architecture choices are worth understanding.Lily runs Qwen3.6-35B-A3B on Apple silicon using a Rust runtime with custom Metal kernels. The optimization split is separate paths for prefill and decode. That matters because the bottleneck on local AI models shifts depending on whether you're processing a long input or generating output tokens. Lily tunes both separately to match what the hardware can actually do.The use case is cleaner…
7:32 PM ET
WAWaymo
7:32 PM ETWAWaymo
@smsehy on X
Waymo designed its own inference ASIC rather than using off-the-shelf GPUs. The thermal envelope inside a vehicle roof module does not allow a 300W GPU. Automotive compute is a packaging and thermal problem first, a silicon problem second. AEC-Q100 qualification on a custom ASIC takes years, but the power budget forces the decision.
3:26 AM ET
@smsehy
3:26 AM ET@smsehy
@smsehy on X
Localized low-latency inference on humanoids requires automotive-grade high-bandwidth memory that withstands intense mechanical vibration and thermal cycling. Sizing on-chip memory buffers to prevent bus saturation during continuous high-rate tactile policy execution directly dictates compute bill of materials.
10:31 PM ET
ZIZiNovaLabs
10:31 PM ETZIZiNovaLabs
LimX Dynamics on X: "Together with ZINOVA's Tool Intelligence, TRON 2 takes on increasingly complex construction workflows.
@ZiNovaLabs builds on TRON 2 to explore an innovative robotic configuration for construction, demonstrating key tasks in a scaled-down tilt-up construction workflow, includi… / X
Post Log in Sign up Post LimX Dynamics on X: "Together with ZINOVA's Tool Intelligence, TRON 2 takes on increasingly complex construction workflows. @ZiNovaLabs builds on TRON 2 to explore an innovative robotic configuration for construction, demonstrating key tasks in a scaled-down tilt-up construction workflow, including formwork assembly, multi-layer rebar placement and tying. TRON 2 serves as a modular and extensible embodied robotic platform for multi-tool, multi-step tasks across large workspaces. Its dual arms handle construction tools and materials across different orientations and…
7:00 PM ET
Nvidia
$NVDA
7:00 PM ETNvidia
Noble Machines Accelerates Humanoid Robot Development 3X | NVIDIA
Manufacturing \| Robotics Noble Machines Accelerates General Purpose Industrial Robot Development With NVIDIA Isaac GR00T Learn More Objective Noble Machines is advancing industrial robot development with the Moby general-purpose industrial robot and its supporting software stack, including a proprietary AI driven whole-body-control system. Moby is designed for complex, hazardous, and labor-intensive work across factories, logistics centers, construction sites, and semiconductor facilities, where operating reliably requires advanced perception, reasoning, adaptation, and physical interaction.…
2:09 PM ET
Nvidia
$NVDA
2:09 PM ETNvidia
Seeed reBot Arm B601-RS: Physical AI & VLA Model Course | NVIDIA DLI Series
Warehouse China Warehouse US Warehouse Germany Warehouse The store will not work correctly in the case when cookies are disabled. US Warehouse: Enjoy FREE UNIUNI shipping on orders under $50! (Excludes XIAO & Raspberry Pi series) \\ \\ the AI Hardware Partner For Industry Products Local Warehouse Documents Customization Solution Software Support Open Claw Quick Order Account Get more with a SeeedStudio accountclear \\ \\ 1 on 1 Support \\ \\ Local Delivery \\ \\ 30-Day DOA Guarantee \\ \\ 30-Day Returns Sign In Create Account My Orders\\ Track, change, cancel Account Information\\ Update…
12:00 PM ET
WAwafer
12:00 PM ETWAwafer
wafer on X: "We’re excited to announce that we’ve raised a $40M Series A!
co-led by @MarathonMP and @chemistry, with participation from @Wing_VC, @AMD Ventures, @outsetcap, @fiftyyears, and @ycombinator, and our existing investors doubling down on @wafer_ai. we are also joined by an incredible lis… / X
Post Log in Sign up Post wafer on X: "We’re excited to announce that we’ve raised a $40M Series A! co-led by @MarathonMP and @chemistry, with participation from @Wing\VC, @AMD Ventures, @outsetcap, @fiftyyears, and @ycombinator, and our existing investors doubling down on @wafer\ai. we are also joined by an incredible list of angels, including @JeffDean (CEO, @DiscoLoopAI), @rauchg (CEO, @vercel), @andyfang (CTO, @DoorDash), @kvogt (CEO, Bot), @akothari (COO, @NotionHQ), @eastdakota (CEO, @Cloudflare), @deepgramscott (CEO, @DeepgramAI), and more! Most inference optimization today is manual,…
10:31 PM ET
Google
$GOOGL
10:31 PM ETGoogle
smartfitguide on X: "Google DeepMind's RT-2 model treats robot actions as text tokens. By co-fine-tuning pre-trained Vision-Language Models with robotic data, it achieved nearly double the performance on novel tasks compared to its predecessor, RT-1. #AI" / X
Post Log in Sign up Post smartfitguide on X: "Google DeepMind's RT-2 model treats robot actions as text tokens. By co-fine-tuning pre-trained Vision-Language Models with robotic data, it achieved nearly double the performance on novel tasks compared to its predecessor, RT-1. \#AI" - smartfitguide @bareani21645 Google DeepMind's RT-2 model treats robot actions as text tokens. By co-fine-tuning pre-trained Vision-Language Models with robotic data, it achieved nearly double the performance on novel tasks compared to its predecessor, RT-1. #AI View media 10:31 PM · Aug 31, 20268Views Reply 1
9:40 PM ET
ZHZhipu
9:40 PM ETZHZhipu
tphuang on X: "GLM-6 disclosed - more parameter & lower activation ratio & inference cost & improved post training expected. Model to make self training decisions
Expecting significant increased domestic compute addition in next 3 to 6 months. Using self owned, leased & purchased c… / X
Post Log in Sign up Post tphuang on X: "GLM-6 disclosed - more parameter & lower activation ratio & inference cost & improved post training expected. Model to make self training decisions Expecting significant increased domestic compute addition in next 3 to 6 months. Using self owned, leased & purchased compute." - tphuang @tphuang Sep 1 Regardless of how Zhipu employees view AGI internally, it publicly follows the OpenAI/Ant line of scaling RSI & Autonomous AI as the future. AI path according to Zhipu: Chat -> Coding -> Co-Work -> Autonomous Larger TAM @ each step w/ last…
9:40 PM ET
@tphuang
9:40 PM ET@tphuang
@tphuang on X
It already went oversea w/ Malaysia deployment & will have more oversea CSP cooperation deal signed in next month or 2. It already has 100k domestic AI chip inference cluster & plan to increase significantly over next 6 months. It did 62T tokens in 6 days w/ Ox Alpha & inference cost has dropped 80% since start of yr
9:40 PM ET
ZHZhipu
9:40 PM ETZHZhipu
tphuang on X: "How it views its own models in terms of the kind of tasks them can do. GLM-5.3 in terms of long time horizon work is still short of 1 wk, need at least 1 month to be autonomous.
Flash is the 1st one to use domestic chips & have low cost + intelligence to support long tasks." / X
Post Log in Sign up Post tphuang on X: "How it views its own models in terms of the kind of tasks them can do. GLM-5.3 in terms of long time horizon work is still short of 1 wk, need at least 1 month to be autonomous. Flash is the 1st one to use domestic chips & have low cost + intelligence to support long tasks." - tphuang @tphuang Sep 1 Regardless of how Zhipu employees view AGI internally, it publicly follows the OpenAI/Ant line of scaling RSI & Autonomous AI as the future. AI path according to Zhipu: Chat -> Coding -> Co-Work -> Autonomous Larger TAM @ each step w/ last step having…
8:36 PM ET
Skild AI
8:36 PM ETSkild AI
The Robot Report: Skild AI unveils S1 robot foundation model | AI Understanding
Back to News ProductAI Understanding briefing The Robot Report: Skild AI unveils S1 robot foundation model The Robot Report reports that Skild AI unveiled S1, a robot foundation model that the company says can learn complex tasks from a single human demonstration video and operate across multiple robot forms. By AI Understanding EditorialSeptember 1, 2026 at 12:36 AM UTCUpdatedSeptember 1, 2026 at 1:47 AM UTC6 min read Read the primary source The short version The Robot Report reports that Skild AI unveiled S1, a robot foundation model that the company says can learn complex tasks from a…
8:00 PM ET
Google
$GOOGL
8:00 PM ETGoogle
Introducing Agentic Video in Gemini
Introducing agentic video understanding with Gemini Sep 01, 2026 \| 7 min read - x.com - Facebook - LinkedIn - Mail - Copy link Our new agentic feature for video analysis cuts token consumption by up to 88%, reduces costs by up to 66%, and boosts quality by up to 7%. Rohan Doshi Senior Product Manager, Google DeepMind Mario Lučić Research Director, Google DeepMind Share - x.com - Facebook - LinkedIn - Mail - Copy link Your browser does not support the audio element. Listen to article \[\[duration\]\] minutes This content is generated by Google AI. Generative AI is experimental…
7:00 PM ET
OPOpenAI
7:00 PM ETOPOpenAI
@SemiAnalysis_ on X
Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) This week Bryan, Myron and Jordan (@JordanNanos) discuss our recent article on OpenAI Jalapeño. They cover the performance, architecture, programming model, implications for NVIDIA and more. 0:00 Cold Open 1:05 Jalapeno Overview 4:22 Tokens Per Megawatt 9:15 Benchmark Caveats 13:48 The CUDA Moat 21:12 How OpenAI Did It 32:13 Samsung HBM4 42:10 AI-Designed Silicon 49:24 Architecture Deep Dive 56:58 Doom and Wrap
11:34 AM ET
Nvidia
$NVDA
11:34 AM ETNvidia
@ParadisLabs on X
Some AI/semis notes from the past week: -> $NVDA huge Q2 revenue and profit growth. 2027+ is more growth even under a bottlenecked environment. Nvidia GOAT company? -> $NBIS first cloud customer for $NVDA Groq 3 LPX inference rack which now entered full production = fuels ongoing narrative that Nebius are Nvidia's fav neocloud. -> SaaS is alive with $CRM proving all the "SaaSpocalypse" doomers wrong. Anthropic announced Claudeforce too = Salesforce / enterprise SaaS data too good. -> Anthropic S-1 coming anyday now. Jensen said investing in Anthropic/OAI is a "once-in-a-generation…
Error 401 (Bad Request)!!1 401. That’s an error. The server cannot process the request because it is malformed. It should not be retried. That’s all we know. Back Skip navigation Search Search with your voice Sign in Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) Tap to unmute 2x Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) SemiAnalysis 19,743 views 4 days ago Copy link Info Shopping If playback doesn't begin shortly, try restarting your device. • You're signed out Videos you watch may be added to the TV's watch history and influence TV…
3:27 PM ET
@frontrunvc
3:27 PM ET@frontrunvc
@frontrunvc on X
plugged the Machine Age memo into claude + the @frontrunvc mcp. companies flagged in the last 30 days that fit the thesis: 👇 chips + compute: @stanmachines AI that designs chips, YC S26 @neurophos a rack in the size and power draw of one GPU @lamblabs custom silicon for inference, 20k tok/s, YC S26 @xlight_inc the world's most powerful lasers @lumilens_ photonic interconnects for AI compute data centers + power: @pacific_ycs26 modular data centers for the hardest environments, YC S26 @teraplex_usa prefab data centers, GB300 ready @vairehq near-zero energy computing @actinideinc unlocking the…
2:10 PM ET
https://x.com/LerrelPinto
2:10 PM EThttps://x.com/LerrelPinto
Lerrel Pinto on X: "Turns out that doing In-context learning for robots is not that hard..." / X
Post Log in Sign up Post Lerrel Pinto on X: "Turns out that doing In-context learning for robots is not that hard..." - Lerrel Pinto @LerrelPinto Turns out that doing In-context learning for robots is not that hard... 00:00 View media 2:10 PM · Aug 28, 2026267.7KViews 26 33 569 205 - ismaelvega @notismaelvega Aug 28 it's right behind me isn't it View media Reply 40 4.9K - Chenhao Li @breadli428 Aug 28 I feel this is a not very well-defined area where we know how “OOD” the task is. 2 22 3.4K - Heeger @GChongkai Aug 28 What's the difference between in-context learning and those one-shot…
11:19 AM ET
ARAxis Robotics
11:19 AM ETARAxis Robotics
Cicada Market Making on X: "https://t.co/26k9SK56mf" / X
Post Log in Sign up Post - Cicada Market Making @cicada\mm Inside Robotics and Physical AI Featuring insights from Axis Robotics In 2026, the artificial intelligence industry learned to solve the compute problem. GPUs are becoming more accessible, models are cheaper to run inference on, cloud infrastructure keeps growing. But the next wave of AI, robotics and Physical AI, has a completely different problem. It all comes down to data that simply does not exist in the volume needed. No Internet for Robots LLMs grew out of a foundation that already existed. Decades of text on the internet,…
6:09 AM ET
Skild
6:09 AM ETSkild
@0xconglomerate on X
Is it really possible for a robot to learn a new task from one human video, with zero fine-tuning? I’ve read a lot of hype around Skild’s S1, with some already calling it robotics’ “ChatGPT moment.” That sounded a bit absurd to me, so I looked into what’s actually going on 👇 ➦ What is @SkildAI actually claiming? I’ve written enough marketing material for AI, robotics, and tech companies that I’m usually skeptical when something sounds too good to be true. If you only saw the announcement post, you could easily think the robot just watches one human video and somehow learns the task from that…
1:19 PM ET
Skild AI
1:19 PM ETSkild AI
Skild AI on X: "Introducing S1, our new foundation model that learns from one example.
It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning.
Watch S1 operate in real-time via in-context learning:" / X
Post Log in Sign up Post Skild AI on X: "Introducing S1, our new foundation model that learns from one example. It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning. Watch S1 operate in real-time via in-context learning:" - Skild AI @SkildAI Introducing S1, our new foundation model that learns from one example. It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning. Watch S1 operate in real-time via in-context learning: 00:00 View media 1:19 PM · Aug 25, 20263.4MViews 427…
11:07 AM ET
NVIDIA
$NVDA
11:07 AM ETNVIDIA
NVIDIA Announces Jetson Orin Nano 2 Robotics Computer to Redefine Entry-Level Edge AI | NVIDIA Newsroom
NVIDIA Announces Jetson Orin Nano 2 Robotics Computer to Redefine Entry-Level Edge AI August 25, 2026 News Summary: - NVIDIA Jetson Orin Nano 2, a new robotics computer for entry-level edge AI, enables millions of developers worldwide to build robots, delivery and inspection drones, and vision AI systems for frontier physical AI applications. - Jetson Orin Nano 2 delivers 2x the inference performance of its predecessor in the same form factor and consumes 40% less power at the same performance. - More than 3 million developers have built on the NVIDIA robotics stack, with Cognex, Doosan…
12:00 PM ET
https://x.com/Majumdar_Ani
12:00 PM EThttps://x.com/Majumdar_Ani
Anirudha Majumdar on X: "Very cool work, Pete! I'm curious if you've tried an extreme version of this: no in-context example at all; just place objects in front of the robot and see if it can infer the task.
I suspect this will have non-trivial success rates, which would also allow you to figure … / X
Post Log in Sign up Post Anirudha Majumdar on X: "Very cool work, Pete! I'm curious if you've tried an extreme version of this: no in-context example at all; just place objects in front of the robot and see if it can infer the task. I suspect this will have non-trivial success rates, which would also allow you to figure out how much of the benefit is actually coming from the prompt (rather than pure zero-shot task inference + capability). I remember Russ showed something like this in his Stanford talk on LBMs last year: https://t.co/0hEBrVlfOQ" - Pete Florence @peteflorence Aug 19 Ever since…
12:00 PM ET
NVIDIA
$NVDA
12:00 PM ETNVIDIA
Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control | NVIDIA Technical Blog
Technical Blog Subscribe Related Resources Robotics Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control Aug 19, 2026 By Saeed Babamohamadi +12 Like Discuss (0) - L - T - F - R - E AI-Generated Summary - NVIDIA Jetson Thor can run the 4B Cosmos 3 Edge omni-model natively for on-device robot manipulation policies. - Post-training Cosmos 3 Edge on the Cosmos3-DROID dataset produces a policy that generates action chunks in about 1.53 seconds on Jetson AGX Thor T5000, enabling continuous real-time control. - In closed-loop RoboLab evaluation, the post-trained Edge policy reaches 22.9%…
NVIDIA12OpenAI6Skild AI4agentic video understanding3Google3tokens processed3AI chip inference cluster2AMD2Anthropic2MiniMax2Samsung2Vercel2Zhipu2a16z1a16z Machine Age Fund.1ADLINK1Agibot1Apple1Astribot1Axis Robotics1BABA1BABA funding for AI.1Booster Robotics1Broadcom1chemistry1Cloudflare1Cognex1Compute requirements for humanoids.1Current hours of available data for training robots.1Current throughput of valid trajectories collected per hour.1d-Matrix1Deep Robotics1DeepgramAI1development timeline of Moby.1development timeline reduction.1DiscoLoopAI1domestic AI chip inference cluster1DoorDash1Doosan Bobcat1Eptolink1fal1Fcc1Fetch Robotics1fiftyyears1Flash1Geely Auto1GEN-1.5 can learn new tasks in a few seconds.1General Catalyst1Groq1H3 Max Turbo1HBM prices.1Hugging Face1inference cost reduction1Innolight1LG Innotek1LimX Dynamics1LM Studio1Lotus Cars1Manycore Tech1MarathonMP1market size for Autonomous AI1Matic1Meta1Moby deployment in a semiconductor facility.1MRVL1Nebius1Noble Machines1NotionHQ1Nscale1Nscale IPO.1Number of global contributors and collected trajectories on the platform.1Ollama1outsetcap1performance metrics.1Perplexity1power consumption comparison of Jetson Orin Nano 21Rockchip1Runway1Salesforce1SambaNova1Samsung Electro Mechanics1Schaeffler1sec1Seeed1Simulated robot bodies1SK Hynix1SKHY Indiana packaging plant.1Solomon1success rate of the post-trained policy.1Target valid hours of data per month from the Mobile Ego-Centric App.1tasks available in RoboLab.1Tesla1Training Data1training needed for current VLA models to match S1's accuracy1TRON 21Unitree1validated run uses 64 nodes of 4× GB200 for 60K iterations, roughly 68 hours (17.4K GB200-hours).1Views on the post1vLLM1wafer1Wafer's funding announcement.1Waymo1Wing1Wing_VC1Yangtze OF1ycombinator1YMTC1ZiNovaLabs1Zite1
H3 Max Turbo$0.01“a $0.01/s video generation endpoint makes AI video economically viable for things like product demo generation, marketing content pipelines, and real-time synthetic data.”@stretchcloud on X
Compute requirements for humanoids.0 units“Sizing on-chip memory buffers to prevent bus saturation during continuous high-rate tactile policy execution directly dictates compute bill of materials.”@smsehy on X
domestic AI chip inference cluster100,000 units“It already has 100k domestic AI chip inference cluster & plan to increase significantly over next 6 months.”@tphuang on X
tokens processed62,000 units“It did 62T tokens in 6 days w/ Ox Alpha”@tphuang on X
inference cost reduction$80“inference cost has dropped 80% since start of yr”@tphuang on X
BABA funding for AI.$10.2B“$BABA selling $10.2B of stock to fund AI.”@ParadisLabs on X
a16z Machine Age Fund.$1.1B“a16z launched $1.1B Machine Age Fund for chips, memory, networking, data centres, cooling, robotics and edge.”@ParadisLabs on X
Nscale IPO.3 units“Nscale (backed by $NVDA) targeting up to $3B in US IPO soon.”@ParadisLabs on X
HBM prices.50 hrs“HBM contract prices maybe +50% or more in 2027 on the back of the $NVDA server hike = memory makers keep printing earnings.”@ParadisLabs on X
SKHY Indiana packaging plant.458 units“$458M CHIPS funding done.”@ParadisLabs on X
Number of global contributors and collected trajectories on the platform.150,000 units“As of late August 2026, the platform has surpassed 150,000 global contributors and collected over 3.7 million trajectories across 4,000+ published tasks.”Cicada Market Making on X: "https://t.co/26k9SK56mf" / X
Simulated robot bodies100,000 units“They trained locomotion across 100,000 simulated robot bodies, which forces the policy to deal with different morphologies instead of just memorizing how one specific robot moves.”@0xconglomerate on X
Reef is one of the most ambitious open-source AI infrastructure projects I have seen this year. ~300 GitHub stars in two days. The core idea is genuinely different from anything else in the space: most RL post-training pipelines treat the model and the scaffolding around it as separate problems. You train one, then you write the other. Reef rejects that separation entirely. It co-evolves model weights and agent harness simultaneously, using live task outcomes as the feedback signal for both. Model side: SAO (Self-Aligned Optimization) adjusts weights from what the agent actually accomplished.…
12:45 AM ET
NVIDIA
$NVDA
12:45 AM ETNVIDIA
@stretchcloud on X
What NVIDIA shipped at IFA yesterday quietly changes local inference infrastructure. PAIR (Personal AI Router) is a free, open-source tool that auto-discovers compatible GPUs on your local network and routes inference requests to whichever machine has capacity. RTX 20-series and above, Apple M4, DGX Spark all supported. Integrates natively with Ollama and LM Studio. A few things to unpack. This is load balancing, not memory pooling. PAIR does not shard a 70B model across three machines or aggregate VRAM. Each machine runs models that fit in its own memory. What PAIR does: intelligently…
12:39 AM ET
Meta
12:39 AM ETMeta
@tchsignal on X
Meta Puts a Price on AI Interaction Data Meta is offering a striking trade-off for developers using its Muse Spark AI model: much cheaper API access if they opt in to Contributor terms that allow their prompts and model outputs to contribute to future model development. Standard pricing is $1.25 per million input tokens and $4.25 per million output tokens. Under Contributor pricing, those rates fall to $0.10 and $0.20 respectively. That distinction matters. Meta is not paying users cash, and the Contributor terms are not the default for every Muse Spark customer. Standard pricing remains…
12:30 AM ET
AMD
12:30 AM ETAMD
@SemiAnalysis_ on X
On total tokens per $ TCO, a new AMD MI355x submission beats B300 at lower interactivity ranges on AgentX. Both high throughput configs use disaggregated setups. Shoutout to vLLM, AMD, and LMCache engineers. https://t.co/Sc35IXkXY8
12:15 AM ET
BRBroadcom
12:15 AM ETBRBroadcom
@ricci_nov on X
Broadcom's core switching GM Asad Khamisy told Semicon Taiwan that co-packaged optics is not necessarily the best answer for scale-up networking, and fits scale-out better. The optics story splits by tier rather than replacing copper everywhere at once. https://t.co/5rPTq1LFXO
11:35 PM ET
Tesla
$TSLA
11:35 PM ETTesla
@elonmusk on X
Cybercab safety https://t.co/9DQYsSTGeU
11:13 PM ET
@smsehy
11:13 PM ET@smsehy
@smsehy on X
A 15-gigawatt deficit in grid power will accelerate the shift toward embedded edge intelligence and model quantization. Running efficient inference directly inside vehicle gateways and factory edge nodes bypasses utility interconnect queues. https://t.co/m8Mi9dmI5E
11:11 PM ET
@smsehy
11:11 PM ET@smsehy
@smsehy on X
Physical constraints across high-bandwidth memory packaging and behind-the-meter power interconnects are forcing compute optimization down to the edge. When centralized megawatts take four years to commission, running quantized models on ruggedized plant controllers is the only operational option.
10:47 PM ET
@stretchcloud
10:47 PM ET@stretchcloud
@stretchcloud on X
The 30-50x cost difference between computer use and MCP keeps rattling around in my head. The framing that clicks for me: computer use is paying for a human-in-the-loop at model prices. Every screen interaction generates a screenshot, feeds it through vision, waits for coordinate output, executes the click, takes another screenshot. You are spending tokens on what amounts to OCR and cursor navigation, round after round. MCP replaces that entire loop with a typed function call. No screenshots. No vision pass. No coordinate uncertainty. The model calls a tool, the tool returns structured data,…
10:27 PM ET
@stretchcloud
10:27 PM ET@stretchcloud
@stretchcloud on X
The thing I keep noticing across agentic frameworks in 2026: resumability is now the reliability primitive everyone is building toward. Genkit Go 1.13 ships it properly. A Generate call that fails at tool round five returns what it finished alongside the error. You pass resp.History() back in, only the failed step reruns. No wasted tool calls. No redone work. The same logic extends to full agent sessions. A failed or cancelled turn saves completed rounds as a snapshot. Send an empty input, it picks up from there. This matters for production. Most AI agent failures today are partial. The agent…
10:14 PM ET
OPOpenAI
10:14 PM ETOPOpenAI
@DrJimFan on X
Good old days at OpenAI in 2016: an agent stares at screen pixels, moves a mouse, and books a flight on United. We called it World of Bits, inside OpenAI Universe. 10 yrs later, Astra is reincarnated in the same universe. Even the naming is astronomically correct 😆 Universe was perhaps the most ambitious AI infra project at the time, but we couldn't quite figure out how to solve it. A policy with zero prior knowledge of what a "submit" button does has to rediscover the entire internet visual lingua by trial and error. In retrospect, RL from scratch against hand-drawn, per-task "artisan"…
5:42 PM ET
OPOpenAI
5:42 PM ETOPOpenAI
@rimtoln on X
GPT-6 ASTRA ISN'T A PATCH. IT'S A GENERATION FLIP. openai just shipped the model they call a generational leap past gpt-5.6 sol not better chat better computer use · coding · cyber · science ▹ why this is the breakthrough first openai model at Critical cyber threshold agent stacks that actually drive the machine browser · forms · repos · multi-step work enterprise-first rollout · daybreak defenders first ▹ the scoreboard (vendor table) frontiermath t4 · astra 97.6 · fable 5.1 87.8 sol was 83.0 · that's a real cliff terminal-bench science · 64.6 vs fable 52.6 automationbench · 41.4 vs 31.4…
3:30 PM ET
@SemiAnalysis_
3:30 PM ET@SemiAnalysis_
@SemiAnalysis_ on X
Shoutout to the cracked team at @vllm_project that implemented recent agentic workload optimizations. (1/5)🧵 https://t.co/BsXCtLBwvU
3:03 PM ET
ASAstribot
3:03 PM ETASAstribot
@XRoboHub on X
Robots can’t stop while the model thinks. In a high-speed throw, even one pause can kill the momentum. Astribot released SmoothRL for online RL during async inference. S1 keeps moving as the model computes the next action chunk. Most actions in a chunk never execute. Train on the full chunk, and RL credits or blames moves that never happened. SmoothRL learns only from executed actions, matching real deployment timing. After 250 rollouts, tossing jumped 39%→94%, pen capping 8%→83%, and box opening 30%→90%. One autonomous toss cut acceleration RMS by 52% and jerk RMS by 47%. The model keeps…
2:25 PM ET
VEVercel
2:25 PM ETVEVercel
@stretchcloud on X
One command now routes Claude Code, Cursor, Codex, Cline, and four other coding harnesses through a single AI Gateway. That is what Vercel's setup flow does: run `vercel ai-gateway coding-agents setup`, it detects which agents you have installed, writes their configuration, and provisions an API key. No markup added on top. 300-plus models from 30-plus providers, behind one endpoint. The practical gain is real. When Anthropic, OpenAI, and SpaceXAI all hit outages in the same window this week, a gateway with automatic fallbacks made the difference between an agent that stalled and one that…
1:43 PM ET
@rimtoln
1:43 PM ET@rimtoln
@rimtoln on X
EVERYONE SHOWS THE 300-AGENT SWARM. almost nobody shows what happens after 300 agents are easy to spawn turning 300 outputs into one structure is the hard part ▹ the pipeline nobody demos collect → connect → organize → unify one context graph at the end the swarm starts as hundreds of isolated research paths then sources start sharing entities links form · contradictions surface weak claims stay exposed eventually the mess collapses into one structure the whole system can reason over ▹ the reframe that's the part you usually never see the swarm is the demo the graph is the product bookmark…
1:32 PM ET
ZIZite
1:32 PM ETZIZite
@stretchcloud on X
The gap between "AI can help you build this" and "AI can actually deploy this for you" just got smaller. Zite released an MCP server that lets Claude, ChatGPT, and Cursor generate fully deployable business applications, not just code files you still have to wire up yourself. The model describes the app, Zite handles the infrastructure and deployment. This is a different category from code generation. Code generation gives you an artifact you own and maintain. Deployment-as-a-tool gives you a running application where the lifecycle management is abstracted away. The distinction matters for the…
1:14 PM ET
Nvidia
$NVDA
1:14 PM ETNvidia
@tchsignal on X
Nvidia RTX Spark N1X Brings Up to 128GB of Unified Memory to Windows PCs Nvidia's RTX Spark N1X is more than another increase in CPU and GPU core counts. The new platform combines a Grace CPU, Blackwell RTX GPU and unified memory in Windows 11 laptops and small desktops. The higher-end configuration pairs 20 CPU cores with 6,144 CUDA cores and supports up to 128GB of LPDDR5X unified memory. A second configuration combines 18 CPU cores with 5,120 CUDA cores and supports up to 64GB according to Nvidia's current specifications. Laptop versions operate within a 45–80W TDP range, while Nvidia…
1:04 PM ET
Nvidia
$NVDA
1:04 PM ETNvidia
@tchsignal on X
Nvidia PAIR Turns Multiple Computers Into a Local AI Compute Pool Nvidia has launched the beta of Personal AI Router, or PAIR, a free software tool designed to make multiple compatible computers on the same local network available for AI inference through a single endpoint. PAIR can route independent requests to available machines based on factors including model availability and GPU utilization. It supports local inference backends including Ollama and LM Studio, with compatible hardware spanning GeForce RTX 20 Series and newer GPUs, RTX PRO systems, DGX Spark and Apple M4 or newer devices.…
12:36 PM ET
Nvidia
$NVDA
12:36 PM ETNvidia
@tchsignal on X
Nvidia’s $12.9B Hugging Face Deal Raises a Bigger Question About Open AI Nvidia has agreed to acquire Hugging Face in a deal valued at about $12.93 billion, but the transaction has not closed yet. It is expected to complete in the first half of 2027, subject to regulatory approvals and other closing conditions. The bigger issue is what Nvidia ownership could mean for an AI platform used by more than 18 million developers, researchers and creators, with over 3 million models, 500,000 datasets and 1 million applications. Nvidia says Hugging Face will remain an open platform. Developers will not…
12:25 PM ET
@MelvinInvests
12:25 PM ET@MelvinInvests
@MelvinInvests on X
AI’s biggest bottleneck is moving data and that could still create huge opportunities for optical networking companies (Save this) The chart shows a 1.6T optical transceiver, a device that transfers data between AI servers, switches, GPUs, and fiber optic networks. 1.6T means it can theoretically move up to 1.6 terabits of data per second, or 1,600 gigabits and that is about twice the speed of an 800G connection. This technology is important because AI data centers contain thousands of GPUs that must constantly exchange information. As AI models become larger, slow connections can leave…
8:48 AM ET
MIMiniMax
8:48 AM ETMIMiniMax
@stretchcloud on X
Video generation just crossed a threshold I didn't expect until 2027. https://t.co/wgZmOqI59X took MiniMax's H3 model, post-trained it for quality and cost, then ran it through a custom inference engine. The result: a 5-second 768p clip renders in about 3 seconds. Faster than real-time. You finish reading the prompt before the video finishes generating. The inference numbers matter because they change the workflow. When generation takes 60+ seconds, you commit to a prompt and wait. When it takes 3 seconds, you iterate. Sketch, reject, sketch again. That shift is what moves video from…
8:46 AM ET
@smsehy
8:46 AM ET@smsehy
@smsehy on X
An autonomous vehicle generates 1 to 2 TB of sensor data per hour. Streaming that to the cloud for inference costs more than the compute hardware in the vehicle. The economics of edge AI in automotive were settled before the debate started.
5:17 AM ET
@smsehy
5:17 AM ET@smsehy
@smsehy on X
Enterprise customers deploying autonomous physical fleets cannot rely on centralized multi-tenant cloud APIs. Full operational sovereignty requires edge inference execution completely decoupled from public network connectivity. https://t.co/YJl2DcWQOa
4:56 AM ET
@smsehy
4:56 AM ET@smsehy
@smsehy on X
In automotive zonal compute, reserving leading-edge nodes for central AI inference while keeping I/O and power interfaces on mature 28nm silicon is the only viable path to maintain AEC-Q100 thermal reliability and control bill of materials cost. https://t.co/a4PfhQqwNR
4:47 AM ET
RORockchip
4:47 AM ETRORockchip
@seeedstudio on X
🔥🔥NEW Launch: reComputer RK3576 Dev Kit/Module/IO Board is Officially HERE!!!! Level up your edge AI development with Rockchip #RK3576 today! 🎉 🧩🧩Dev Kit / Module / IO Board: https://t.co/xQVIT62cbw Powered by #Rockchip #RK3576 octa-core #CPU + 6 TOPS #NPU, this series delivers reliable, low-latency on-device AI inference without cloud dependency, fully empowering industrial edge intelligence, smart vision, local #LLM/#VLM, and offline speech AI applications. 💻🤖 ✅ Powerful Edge AI Performance -YOLO11 at 77.9FPS (640×640) -DeepSeek/Qwen/Whisper -8K video decoding and 4K@60fps encoding…
4:14 AM ET
@OperationsPLS
4:14 AM ET@OperationsPLS
@OperationsPLS on X
Edge AI + Physical AI is the architecture shift. Moving inference to the robot eliminates cloud latency — critical for real-time manipulation and safety. HBM5 at the edge means robots think at the point of action. https://t.co/Z8PWXWTWjM
10:14 PM ET
@stretchcloud
10:14 PM ET@stretchcloud
@stretchcloud on X
Sodium v2 introduces a WebMCP layer. Any website can now serve structured data directly to an AI agent without HTML scraping. This is AEO (Agent Engine Optimization) in practice. The web is starting to build for agents, not just browsers. I built DeepScrape to handle exactly this: giving agents reliable, structured access to web data at inference time. https://t.co/JZYTWY8roq https://t.co/oUD6R0G9j7
8:34 PM ET
GCGeneral Catalyst
8:34 PM ETGCGeneral Catalyst
@stretchcloud on X
12 PhDs and a Fields Medal winner built this in 4 months, backed by General Catalyst. https://t.co/pyhyV4u6I4 has a protocol that passes hidden states directly from a 753B frontier model to a 4B edge model at inference time. No text exchanged between the models. Neither model is fine-tuned. They come from different model families entirely. The result: 80% as accurate as the frontier model alone, running 20x faster. First place on ARC-AGI. The mechanism is what matters here. Most model compression strategies force a choice between accuracy and cost. You either run the big model and accept high…
8:14 PM ET
MIMiniMax
8:14 PM ETMIMiniMax
@stretchcloud on X
The pattern in AI inference is consistent: the fast-and-cheap version of a frontier model ships 2-4 weeks after the frontier version. H3 Max Turbo confirms this for video generation.H3 Max launched on August 26 at $0.08 per second of 768p output. H3 Max Turbo launches today at $0.01 per second, runs at 2x the speed, and hits the 97th percentile of H3 Max quality on evals. Seven days to a 50% cost reduction and 2x throughput improvement.The mechanism matters: fal built a separate inference engine specifically around MiniMax's H3 open-weights model architecture. They're not running H3 Max…
8:13 PM ET
@ricci_nov
8:13 PM ET@ricci_nov
@ricci_nov on X
HBQ: W4A5 matching W4A16 accuracy, in less silicon area than NVFP4. On a 28nm ASIC it shows 2.3x area and 4.6x energy efficiency over weight-only quantization, and 1.6-3.3x lower system energy than prior block quantization. Inference cost is still a datapath problem. https://t.co/08u6ZPO2mO
7:34 PM ET
PEPerplexity
7:34 PM ETPEPerplexity
@stretchcloud on X
The hybrid compute pattern is getting serious. Perplexity just open-sourced Lily, the local inference engine powering their Mac hybrid compute feature, and the architecture choices are worth understanding.Lily runs Qwen3.6-35B-A3B on Apple silicon using a Rust runtime with custom Metal kernels. The optimization split is separate paths for prefill and decode. That matters because the bottleneck on local AI models shifts depending on whether you're processing a long input or generating output tokens. Lily tunes both separately to match what the hardware can actually do.The use case is cleaner…
7:32 PM ET
WAWaymo
7:32 PM ETWAWaymo
@smsehy on X
Waymo designed its own inference ASIC rather than using off-the-shelf GPUs. The thermal envelope inside a vehicle roof module does not allow a 300W GPU. Automotive compute is a packaging and thermal problem first, a silicon problem second. AEC-Q100 qualification on a custom ASIC takes years, but the power budget forces the decision.
3:26 AM ET
@smsehy
3:26 AM ET@smsehy
@smsehy on X
Localized low-latency inference on humanoids requires automotive-grade high-bandwidth memory that withstands intense mechanical vibration and thermal cycling. Sizing on-chip memory buffers to prevent bus saturation during continuous high-rate tactile policy execution directly dictates compute bill of materials.
10:31 PM ET
ZIZiNovaLabs
10:31 PM ETZIZiNovaLabs
LimX Dynamics on X: "Together with ZINOVA's Tool Intelligence, TRON 2 takes on increasingly complex construction workflows.
@ZiNovaLabs builds on TRON 2 to explore an innovative robotic configuration for construction, demonstrating key tasks in a scaled-down tilt-up construction workflow, includi… / X
Post Log in Sign up Post LimX Dynamics on X: "Together with ZINOVA's Tool Intelligence, TRON 2 takes on increasingly complex construction workflows. @ZiNovaLabs builds on TRON 2 to explore an innovative robotic configuration for construction, demonstrating key tasks in a scaled-down tilt-up construction workflow, including formwork assembly, multi-layer rebar placement and tying. TRON 2 serves as a modular and extensible embodied robotic platform for multi-tool, multi-step tasks across large workspaces. Its dual arms handle construction tools and materials across different orientations and…
7:00 PM ET
Nvidia
$NVDA
7:00 PM ETNvidia
Noble Machines Accelerates Humanoid Robot Development 3X | NVIDIA
Manufacturing \| Robotics Noble Machines Accelerates General Purpose Industrial Robot Development With NVIDIA Isaac GR00T Learn More Objective Noble Machines is advancing industrial robot development with the Moby general-purpose industrial robot and its supporting software stack, including a proprietary AI driven whole-body-control system. Moby is designed for complex, hazardous, and labor-intensive work across factories, logistics centers, construction sites, and semiconductor facilities, where operating reliably requires advanced perception, reasoning, adaptation, and physical interaction.…
2:09 PM ET
Nvidia
$NVDA
2:09 PM ETNvidia
Seeed reBot Arm B601-RS: Physical AI & VLA Model Course | NVIDIA DLI Series
Warehouse China Warehouse US Warehouse Germany Warehouse The store will not work correctly in the case when cookies are disabled. US Warehouse: Enjoy FREE UNIUNI shipping on orders under $50! (Excludes XIAO & Raspberry Pi series) \\ \\ the AI Hardware Partner For Industry Products Local Warehouse Documents Customization Solution Software Support Open Claw Quick Order Account Get more with a SeeedStudio accountclear \\ \\ 1 on 1 Support \\ \\ Local Delivery \\ \\ 30-Day DOA Guarantee \\ \\ 30-Day Returns Sign In Create Account My Orders\\ Track, change, cancel Account Information\\ Update…
12:00 PM ET
WAwafer
12:00 PM ETWAwafer
wafer on X: "We’re excited to announce that we’ve raised a $40M Series A!
co-led by @MarathonMP and @chemistry, with participation from @Wing_VC, @AMD Ventures, @outsetcap, @fiftyyears, and @ycombinator, and our existing investors doubling down on @wafer_ai. we are also joined by an incredible lis… / X
Post Log in Sign up Post wafer on X: "We’re excited to announce that we’ve raised a $40M Series A! co-led by @MarathonMP and @chemistry, with participation from @Wing\VC, @AMD Ventures, @outsetcap, @fiftyyears, and @ycombinator, and our existing investors doubling down on @wafer\ai. we are also joined by an incredible list of angels, including @JeffDean (CEO, @DiscoLoopAI), @rauchg (CEO, @vercel), @andyfang (CTO, @DoorDash), @kvogt (CEO, Bot), @akothari (COO, @NotionHQ), @eastdakota (CEO, @Cloudflare), @deepgramscott (CEO, @DeepgramAI), and more! Most inference optimization today is manual,…
10:31 PM ET
Google
$GOOGL
10:31 PM ETGoogle
smartfitguide on X: "Google DeepMind's RT-2 model treats robot actions as text tokens. By co-fine-tuning pre-trained Vision-Language Models with robotic data, it achieved nearly double the performance on novel tasks compared to its predecessor, RT-1. #AI" / X
Post Log in Sign up Post smartfitguide on X: "Google DeepMind's RT-2 model treats robot actions as text tokens. By co-fine-tuning pre-trained Vision-Language Models with robotic data, it achieved nearly double the performance on novel tasks compared to its predecessor, RT-1. \#AI" - smartfitguide @bareani21645 Google DeepMind's RT-2 model treats robot actions as text tokens. By co-fine-tuning pre-trained Vision-Language Models with robotic data, it achieved nearly double the performance on novel tasks compared to its predecessor, RT-1. #AI View media 10:31 PM · Aug 31, 20268Views Reply 1
9:40 PM ET
ZHZhipu
9:40 PM ETZHZhipu
tphuang on X: "GLM-6 disclosed - more parameter & lower activation ratio & inference cost & improved post training expected. Model to make self training decisions
Expecting significant increased domestic compute addition in next 3 to 6 months. Using self owned, leased & purchased c… / X
Post Log in Sign up Post tphuang on X: "GLM-6 disclosed - more parameter & lower activation ratio & inference cost & improved post training expected. Model to make self training decisions Expecting significant increased domestic compute addition in next 3 to 6 months. Using self owned, leased & purchased compute." - tphuang @tphuang Sep 1 Regardless of how Zhipu employees view AGI internally, it publicly follows the OpenAI/Ant line of scaling RSI & Autonomous AI as the future. AI path according to Zhipu: Chat -> Coding -> Co-Work -> Autonomous Larger TAM @ each step w/ last…
9:40 PM ET
@tphuang
9:40 PM ET@tphuang
@tphuang on X
It already went oversea w/ Malaysia deployment & will have more oversea CSP cooperation deal signed in next month or 2. It already has 100k domestic AI chip inference cluster & plan to increase significantly over next 6 months. It did 62T tokens in 6 days w/ Ox Alpha & inference cost has dropped 80% since start of yr
9:40 PM ET
ZHZhipu
9:40 PM ETZHZhipu
tphuang on X: "How it views its own models in terms of the kind of tasks them can do. GLM-5.3 in terms of long time horizon work is still short of 1 wk, need at least 1 month to be autonomous.
Flash is the 1st one to use domestic chips & have low cost + intelligence to support long tasks." / X
Post Log in Sign up Post tphuang on X: "How it views its own models in terms of the kind of tasks them can do. GLM-5.3 in terms of long time horizon work is still short of 1 wk, need at least 1 month to be autonomous. Flash is the 1st one to use domestic chips & have low cost + intelligence to support long tasks." - tphuang @tphuang Sep 1 Regardless of how Zhipu employees view AGI internally, it publicly follows the OpenAI/Ant line of scaling RSI & Autonomous AI as the future. AI path according to Zhipu: Chat -> Coding -> Co-Work -> Autonomous Larger TAM @ each step w/ last step having…
8:36 PM ET
Skild AI
8:36 PM ETSkild AI
The Robot Report: Skild AI unveils S1 robot foundation model | AI Understanding
Back to News ProductAI Understanding briefing The Robot Report: Skild AI unveils S1 robot foundation model The Robot Report reports that Skild AI unveiled S1, a robot foundation model that the company says can learn complex tasks from a single human demonstration video and operate across multiple robot forms. By AI Understanding EditorialSeptember 1, 2026 at 12:36 AM UTCUpdatedSeptember 1, 2026 at 1:47 AM UTC6 min read Read the primary source The short version The Robot Report reports that Skild AI unveiled S1, a robot foundation model that the company says can learn complex tasks from a…
8:00 PM ET
Google
$GOOGL
8:00 PM ETGoogle
Introducing Agentic Video in Gemini
Introducing agentic video understanding with Gemini Sep 01, 2026 \| 7 min read - x.com - Facebook - LinkedIn - Mail - Copy link Our new agentic feature for video analysis cuts token consumption by up to 88%, reduces costs by up to 66%, and boosts quality by up to 7%. Rohan Doshi Senior Product Manager, Google DeepMind Mario Lučić Research Director, Google DeepMind Share - x.com - Facebook - LinkedIn - Mail - Copy link Your browser does not support the audio element. Listen to article \[\[duration\]\] minutes This content is generated by Google AI. Generative AI is experimental…
7:00 PM ET
OPOpenAI
7:00 PM ETOPOpenAI
@SemiAnalysis_ on X
Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) This week Bryan, Myron and Jordan (@JordanNanos) discuss our recent article on OpenAI Jalapeño. They cover the performance, architecture, programming model, implications for NVIDIA and more. 0:00 Cold Open 1:05 Jalapeno Overview 4:22 Tokens Per Megawatt 9:15 Benchmark Caveats 13:48 The CUDA Moat 21:12 How OpenAI Did It 32:13 Samsung HBM4 42:10 AI-Designed Silicon 49:24 Architecture Deep Dive 56:58 Doom and Wrap
11:34 AM ET
Nvidia
$NVDA
11:34 AM ETNvidia
@ParadisLabs on X
Some AI/semis notes from the past week: -> $NVDA huge Q2 revenue and profit growth. 2027+ is more growth even under a bottlenecked environment. Nvidia GOAT company? -> $NBIS first cloud customer for $NVDA Groq 3 LPX inference rack which now entered full production = fuels ongoing narrative that Nebius are Nvidia's fav neocloud. -> SaaS is alive with $CRM proving all the "SaaSpocalypse" doomers wrong. Anthropic announced Claudeforce too = Salesforce / enterprise SaaS data too good. -> Anthropic S-1 coming anyday now. Jensen said investing in Anthropic/OAI is a "once-in-a-generation…
Error 401 (Bad Request)!!1 401. That’s an error. The server cannot process the request because it is malformed. It should not be retried. That’s all we know. Back Skip navigation Search Search with your voice Sign in Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) Tap to unmute 2x Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) SemiAnalysis 19,743 views 4 days ago Copy link Info Shopping If playback doesn't begin shortly, try restarting your device. • You're signed out Videos you watch may be added to the TV's watch history and influence TV…
3:27 PM ET
@frontrunvc
3:27 PM ET@frontrunvc
@frontrunvc on X
plugged the Machine Age memo into claude + the @frontrunvc mcp. companies flagged in the last 30 days that fit the thesis: 👇 chips + compute: @stanmachines AI that designs chips, YC S26 @neurophos a rack in the size and power draw of one GPU @lamblabs custom silicon for inference, 20k tok/s, YC S26 @xlight_inc the world's most powerful lasers @lumilens_ photonic interconnects for AI compute data centers + power: @pacific_ycs26 modular data centers for the hardest environments, YC S26 @teraplex_usa prefab data centers, GB300 ready @vairehq near-zero energy computing @actinideinc unlocking the…
2:10 PM ET
https://x.com/LerrelPinto
2:10 PM EThttps://x.com/LerrelPinto
Lerrel Pinto on X: "Turns out that doing In-context learning for robots is not that hard..." / X
Post Log in Sign up Post Lerrel Pinto on X: "Turns out that doing In-context learning for robots is not that hard..." - Lerrel Pinto @LerrelPinto Turns out that doing In-context learning for robots is not that hard... 00:00 View media 2:10 PM · Aug 28, 2026267.7KViews 26 33 569 205 - ismaelvega @notismaelvega Aug 28 it's right behind me isn't it View media Reply 40 4.9K - Chenhao Li @breadli428 Aug 28 I feel this is a not very well-defined area where we know how “OOD” the task is. 2 22 3.4K - Heeger @GChongkai Aug 28 What's the difference between in-context learning and those one-shot…
11:19 AM ET
ARAxis Robotics
11:19 AM ETARAxis Robotics
Cicada Market Making on X: "https://t.co/26k9SK56mf" / X
Post Log in Sign up Post - Cicada Market Making @cicada\mm Inside Robotics and Physical AI Featuring insights from Axis Robotics In 2026, the artificial intelligence industry learned to solve the compute problem. GPUs are becoming more accessible, models are cheaper to run inference on, cloud infrastructure keeps growing. But the next wave of AI, robotics and Physical AI, has a completely different problem. It all comes down to data that simply does not exist in the volume needed. No Internet for Robots LLMs grew out of a foundation that already existed. Decades of text on the internet,…
6:09 AM ET
Skild
6:09 AM ETSkild
@0xconglomerate on X
Is it really possible for a robot to learn a new task from one human video, with zero fine-tuning? I’ve read a lot of hype around Skild’s S1, with some already calling it robotics’ “ChatGPT moment.” That sounded a bit absurd to me, so I looked into what’s actually going on 👇 ➦ What is @SkildAI actually claiming? I’ve written enough marketing material for AI, robotics, and tech companies that I’m usually skeptical when something sounds too good to be true. If you only saw the announcement post, you could easily think the robot just watches one human video and somehow learns the task from that…
1:19 PM ET
Skild AI
1:19 PM ETSkild AI
Skild AI on X: "Introducing S1, our new foundation model that learns from one example.
It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning.
Watch S1 operate in real-time via in-context learning:" / X
Post Log in Sign up Post Skild AI on X: "Introducing S1, our new foundation model that learns from one example. It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning. Watch S1 operate in real-time via in-context learning:" - Skild AI @SkildAI Introducing S1, our new foundation model that learns from one example. It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning. Watch S1 operate in real-time via in-context learning: 00:00 View media 1:19 PM · Aug 25, 20263.4MViews 427…
11:07 AM ET
NVIDIA
$NVDA
11:07 AM ETNVIDIA
NVIDIA Announces Jetson Orin Nano 2 Robotics Computer to Redefine Entry-Level Edge AI | NVIDIA Newsroom
NVIDIA Announces Jetson Orin Nano 2 Robotics Computer to Redefine Entry-Level Edge AI August 25, 2026 News Summary: - NVIDIA Jetson Orin Nano 2, a new robotics computer for entry-level edge AI, enables millions of developers worldwide to build robots, delivery and inspection drones, and vision AI systems for frontier physical AI applications. - Jetson Orin Nano 2 delivers 2x the inference performance of its predecessor in the same form factor and consumes 40% less power at the same performance. - More than 3 million developers have built on the NVIDIA robotics stack, with Cognex, Doosan…
12:00 PM ET
https://x.com/Majumdar_Ani
12:00 PM EThttps://x.com/Majumdar_Ani
Anirudha Majumdar on X: "Very cool work, Pete! I'm curious if you've tried an extreme version of this: no in-context example at all; just place objects in front of the robot and see if it can infer the task.
I suspect this will have non-trivial success rates, which would also allow you to figure … / X
Post Log in Sign up Post Anirudha Majumdar on X: "Very cool work, Pete! I'm curious if you've tried an extreme version of this: no in-context example at all; just place objects in front of the robot and see if it can infer the task. I suspect this will have non-trivial success rates, which would also allow you to figure out how much of the benefit is actually coming from the prompt (rather than pure zero-shot task inference + capability). I remember Russ showed something like this in his Stanford talk on LBMs last year: https://t.co/0hEBrVlfOQ" - Pete Florence @peteflorence Aug 19 Ever since…
12:00 PM ET
NVIDIA
$NVDA
12:00 PM ETNVIDIA
Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control | NVIDIA Technical Blog
Technical Blog Subscribe Related Resources Robotics Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control Aug 19, 2026 By Saeed Babamohamadi +12 Like Discuss (0) - L - T - F - R - E AI-Generated Summary - NVIDIA Jetson Thor can run the 4B Cosmos 3 Edge omni-model natively for on-device robot manipulation policies. - Post-training Cosmos 3 Edge on the Cosmos3-DROID dataset produces a policy that generates action chunks in about 1.53 seconds on Jetson AGX Thor T5000, enabling continuous real-time control. - In closed-loop RoboLab evaluation, the post-trained Edge policy reaches 22.9%…