Raw token throughput is an incomplete way to evaluate AI infrastructure. What decides real value is how that hardware performs against an actual enterprise workload, at an actual cost. Signal65 ran a three-level evaluation of the AMD Instinct MI355X GPU platform against the NVIDIA HGX B200 GPU platform, mapping raw performance and operational economics onto three production workloads. In our testing: ➡️ Financial analysis processed 1.3x more documents per hour at 53% lower cost per document ➡️ Long report generation traded about 5% throughput for 37% lower cost per page ➡️ Interactive RAG chat served the same user population at 39.6% lower cost per hour and 42% lower p99 latency The pattern held across every workload we tested. Cost and workload efficiency prove just as consequential as peak throughput in determining real business value. Full report: https://lnkd.in/eTr3T2Aq
Signal65
Technology, Information and Media
We are here to ensure our partners become the signal of innovation in the noise of the technology markets.
About us
Signal65 provides holistic integrated technical marketing capabilities to our customers. With modern performance and validation labs across datacenter, edge and client segments, we build data-driven tools and assets to showcase material benefits of our customers products and solutions. The experience on HPC, AI, hybrid cloud analysis, storage, and HCI verticals from the team at Futurum Labs merges with a leadership team that brings a background in client solutions, platforms, AI, and graphics from a Fortune 100 compute leader. With a unique mix of industry veterans, engineering minded technicians and writers, and key trusted, independent voices, Signal65 can create and lead a narrative all the way from validation and testing, to messaging and marketing, and even thought leadership and story representation. We are here to ensure our partners become the signal of innovation in the noise of the technology markets.
- Website
-
www.signal65.com
External link for Signal65
- Industry
- Technology, Information and Media
- Company size
- 11-50 employees
- Type
- Privately Held
Employees at Signal65
Updates
-
Peak throughput is the number that often gets the headline but price-performance is the number that shows up on the invoice. Signal65 tested the AMD Instinct MI355X against the NVIDIA HGX B200 on three production models. On tokens per dollar, using blended on-demand cloud pricing, MI355X led on all three in our testing: ➡️ GPT-OSS-120B delivered 2.15x more tokens per dollar ➡️ Qwen3-Next-80B delivered 2.11x more tokens per dollar ➡️ Kimi-K2.6 delivered 1.42x more tokens per dollar On Kimi-K2.6, the B200 posted 17% higher raw throughput, but MI355X still returned more tokens per dollar. Economics and raw performance are different measures, and both belong in an inference sizing decision. Full report: https://lnkd.in/eTr3T2Aq
-
-
Match the Mac, then do more. Beyond the benchmarks, Microsoft Windows 11 covers ground the Mac lineup does not: PC gaming, touch and pen 2-in-1s, displays from 11 to 17 inches and OLED, full-size connectivity, and real choice across price points. Same categories, more capability. Read the report: https://lnkd.in/gjKEPcju
-
-
The assumption that Macs are simply faster no longer holds the way it used to. In head-to-head lab testing, a mainstream Microsoft Windows 11 PC matched the MacBook Air M5 on single-threaded and creative work, and pulled ahead where it counts, in our testing: ✅ +92% multi-thread CPU (Cinebench 2026) ✅ +78% on-device AI (Procyon AI, NPU) ✅ +50% Microsoft Word Full report: https://lnkd.in/gjKEPcju
-
-
When a small team shares one AI appliance, the number that matters is not just single-user speed but how the system holds up when several people query it at once. We ran the NVIDIA DGX Spark as a RAG chatbot on a 30B model across 1, 2, 4, and 8 concurrent users to find out. Per-user throughput dips under load, but total system throughput climbs about 4x as the vLLM continuous batching interleaves the request streams rather than serializing them. Full case study: https://lnkd.in/ehwwME38 What we found in our testing: ➡️ Single-user throughput ran about 56 tokens per second, tapering to about 28.5 tps at 8 users, still above the comfortable reading rate ➡️ Aggregate throughput rose from about 56 tokens per second at 1 user to about 228 tps at 8 users, roughly 4x more total work under load ➡️ 4 concurrent users saw a 7.4s time to first token at about 39 tokens per second each, a responsive experience for interactive chat ➡️ Latency and throughput taper gradually rather than collapsing, so a team can grow into the appliance and sense the limits before they turn disruptive
-
-
Signal65 reposted this
A "200K context window" is a capacity spec. It is not a performance guarantee. In a new benchmark study, Signal65 and Kamiwaza tested 91 models on contamination-resistant retrieval across context sizes from 32K to 200K tokens. The gap between the spec sheet and production reality was hard to miss. At 32K, 27 models held at least 95% retrieval accuracy. At 200K, only three did. Multi-document aggregation, the synthesis work most enterprise use cases actually depend on, degraded even faster. Accuracy fell 26% on average at 200K, and one model dropped from 92% to 24%. A bigger model isn't always the solution; instead, we found that reasoning models, those that deliberate before answering, swept the top spots. In fact, one model's accuracy jumped from 59% to 97% simply by enabling reasoning, proving that how you configure a model can be just as important as the model you select. The practical takeaway for technology leaders: stop asking vendors for a token cap, and start asking for a retrieval-stability curve, broken out by single-document versus aggregation tasks, with reasoning settings made explicit. Peak accuracy on a short context is the easiest number to advertise and the least predictive of production performance. Full breakdown of the RIKER benchmark, and what it means for how you buy and build, is in the blog (link in the comments). Which context size does your current RAG deployment actually hold accuracy at, and how would you know?
-
-
The desktop has powered corporate productivity for decades, yet the AI PC conversation has stayed focused on laptops. Our new study asks whether AI changes the math on the desktop, and whether the silicon underneath the workload actually matters. Full report: https://lnkd.in/enxCZmJU Running a common set of office tasks on an HP EliteDesk 8 SFF with the AMD Ryzen AI 7 PRO 450G, AI assistance cut total time to 19.1 minutes from 56.7 minutes done manually, a 66% reduction in our testing. Across the same testing: ➡️ Individual workflow steps ran as much as 17x faster with AI, with meeting summarization showing the largest gain ➡️ At typical task frequencies, AI users save roughly 1.1 workdays per week, about 11 work weeks per year ➡️ The current-generation chip ran the same AI workloads 1.8x faster than a prior-generation Ryzen desktop ➡️ On some workflows the older desktop ran AI slower than the manual approach, a sign that modern silicon is required to make these workflows worth running at all
-
-
Small teams want production-grade AI without a cloud subscription, a server room, or a dedicated infrastructure hire. We ran the NVIDIA DGX Spark through a full day/night cycle, serving RAG inference for the team by day and fine-tuning the same model with LoRA overnight, to see whether one desktop can cover both halves of a real AI deployment. Full case study: https://lnkd.in/ehwwME38 What the case study found: ➡️ In our testing, 4 concurrent users saw a 7.4s time to first token at roughly 39 tokens per second each, with total system throughput rising about 4x from 1 to 8 users ➡️ In our testing, overnight LoRA fine-tuning held about 425 tokens per second at steady state, finishing a 500-step run in roughly 45 minutes ➡️ 128 GB of unified memory serves a 30B model in NVFP4 by day and fine-tunes the same model with bf16 LoRA at night, with no changes to the model itself ➡️ At $4,700 across 8 users, the hardware works out to under $600 per seat, a one-time cost with no cloud meter and no data leaving the building
-
-
Signal65 reposted this
#Liquid_Cooled_Networking Has Arrived Diana Blass with Signal65 recently got hands-on with Dell Technologies' #PowerSwitch_Z9964FL at #DellTechWorld, our first liquid-cooled data center switch and a major step forward for AI infrastructure. At the heart of this switch is the Broadcom Tomahawk 6 ASIC, purpose-built to move massive amounts of data between GPUs. With 64 ports each delivering 1.6 terabit capacity, the power demands have nearly doubled from the previous generation. #Traditional #air_cooling simply cannot keep up. The Z9964FL uses direct liquid cooling on the #ASIC, the hottest component in the switch. Liquid transfers the heat, carries it away to a Coolant Distribution Unit, and returns in a continuous loop. The result is more efficient cooling that directly improves operational performance, including #AI_token generation throughput. Deploying these liquid cooled racks and solutions often requires plumbing, a CDU, and thoughtful infrastructure planning, but the payoff is significant. Does everything need to be liquid-cooled going forward? No. Dell offers fully air-cooled, fully liquid-cooled, or hybrid deployment models. For high-density AI racks that are difficult to cool with air alone, pairing a liquid-cooled switch alongside them keeps the entire rack clean and efficient. This is part of the Dell Technologies #AI_Factory approach, bringing compute, storage, networking, and cooling together into one architecture. Learn more at Dell.com #DellTechnologies #DellTechWorld #DTW26 #AIFactory #LiquidCooling #DataCenter #Networking #PowerSwitch, #IWork4Dell, James Wynia,
-
Signal65 reposted this
Intel announced new Ethernet Adapters, E835 series at Computex 2026. They highlighted the E835 series' manageability, security features, and competitive power efficiency compared with Nvidia ConnectX-6 Dx and Broadcom P2100G. We also noticed the Precision Time Protocol (PTP) benchmark conducted by Signal65. The results showed that the Intel E835 series had an advantage over Nvidia ConnectX-7 and Broadcom P425G in network timing workloads. https://lnkd.in/g76a_mjk