Skip to content
Morning Briefing · Monday, July 13, 2026

DriveNets Turns Two Data Centers 52 Miles Apart Into One GPU Supercluster

datacenternetwork-automationai-mlnetworkingscience
Listen to the episode
DriveNets Turns Two Data Centers 52 Miles Apart Into One GPU Supercluster
37 min · 168 turns
Plate Irack · row
Datacenter row — per-rack utilization at a glance. Cool colors are slack; warmer fills are pressure.
Top Highlights
№ 01·Top Highlights

Top 3 Highlights

1. DriveNets Turns Two Data Centers 52 Miles Apart Into One GPU Supercluster

TL;DR: DriveNets says it's running two separate GPU sites, 52 miles apart, as a single logical training cluster — a live commercial answer to the exact power-and-land bottleneck Friday's lead story described as a structural constraint on AI buildouts.

Key Points:

  • WhiteFiber's two H200 GPU clusters, 52 miles apart, linked into one logical supercluster via DriveNets' "scale-across" architecture — described in press coverage as the first commercial deployment of this kind
  • Reported specs: 111.2 Tbps of aggregate bandwidth and a guaranteed 0.9 millisecond latency between sites — tight enough that a distributed training job sees one fabric, not two
  • [unverified] We could not independently confirm those exact figures — both the primary HPCwire writeup and DriveNets' own press release returned access errors on direct fetch, so treat the specific numbers as reported by secondary coverage, not confirmed against the source
  • Same week, DriveNets unveiled its next fabric platforms — the 2600SL and 2601S — built on Broadcom's Tomahawk 6 ASIC, 102.4 Tbps per box across 64 ports of 1.6 Tbps each, aimed at 100,000-plus-XPU clusters, shipping Q3 2026
  • Dell'Oro reports Ethernet switch sales inside AI back-end networks more than doubled in Q1 2026 and now make up roughly two-thirds of datacenter switch sales in AI clusters — standard, merchant-silicon Ethernet displacing proprietary interconnect at frontier scale

Deep Dive: This is the most direct technical answer yet to Friday's lead story — AI datacenter opposition groups now number 430, and roughly $130 billion in projects got blocked or delayed in a single quarter. If you can't get the power or the land approved in one place, the obvious move is to stop insisting the whole cluster lives in one place. Scale-across architecture decouples "how big is my GPU cluster" from "how much power can I get approved at a single site" — build wherever power and land actually clear, then bridge sites with a fabric tight enough, and consistent enough, that a training job never notices the seam.

That's a real inversion of a decade of hyperscale networking assumptions. Rail-optimized, tightly-coupled AI fabrics have always leaned on physical proximity — the entire premise of NVLink- and InfiniBand-style interconnects is that the GPUs share a building, often a row of racks. If DriveNets' commercial deployment holds up under scrutiny, "one enormous datacenter" stops being the only path to frontier-scale training, and the binding constraint shifts from site-level power and land to long-haul dark fiber and DWDM availability between sites — a different, and arguably more tractable, problem.

Pair that with the Dell'Oro number: Ethernet, not proprietary interconnect, is now carrying two-thirds of AI back-end network traffic. Scale-across architecture, current-generation Tomahawk silicon, and commodity high-radix Ethernet switching are all pointing the same direction — the AI fabric layer is converging on open, merchant-silicon Ethernet rather than the closed, vendor-specific interconnects that defined the first wave of AI datacenter builds.

If you can't get the power in one place, stop trying to put the cluster in one place — that's the actual argument buried inside a fifty-two-mile fiber run.

So What? If you're anywhere near AI fabric design, start treating long-haul DCI — dark fiber, DWDM, latency budget — as a first-class part of the AI cluster conversation, not a side topic for the WAN team to handle later. Scale-across turns a single-site power approval into a solvable multi-site interconnect problem instead of a hard blocker. Just don't repeat DriveNets' own bandwidth and latency figures as confirmed fact until a primary source or independent third party backs them.

SourcesHPCwire, DriveNets


2. Cisco and Extreme Both Bet the Network Is the Agentic AI Control Plane — Nobody's Answered Who's Accountable

TL;DR: Cisco's new Cloud Control platform and Extreme's Agent One both stake out the same claim this cycle — that the network, not any single AI model, should be where autonomous agents get real operational authority. Simon Willison's essay on "Directly Responsible Individuals," published the same week, is the question neither launch actually answers: who's accountable when the agent gets it wrong.

Key Points:

  • Cisco Cloud Control (unveiled at Cisco Live, global availability targeted July 2026): a unified platform explicitly pitched for "humans and AI agents" to jointly manage, monitor, and defend infrastructure — includes "AI Canvas" telemetry, a "Deep Network Model" trained on forty years of Cisco's own operational data, a natural-language custom-agent builder called Cloud Control Studio, no-reboot runtime vulnerability defense (Live Protect) for Nexus 9000 switches, and quantum-safe secure boot now standard on new routers, switches, and firewalls
  • Extreme Networks shipped Agent One Coworker now, with Agent One Operator due in the fourth quarter, adding a unified "living" topology view and zero-touch provisioning to Platform ONE
  • Cisco's framing: "the network is more powerful than the node" — the pitch is coordination and visibility across systems, not any one model's raw capability; notably light on hard bandwidth or latency benchmarks, which is the tell that this is a control-plane and visibility pitch, not a performance claim
  • The same week, Simon Willison published "Directly Responsible Individuals," arguing accountability is structurally human — you can't fire, discipline, or hold an AI agent legally liable, so handing it real ownership just creates a gap where a human still answers for the outcome with less visibility into how the decision got made. He anchors it on an IBM internal training slide from 1979: "A computer can never be held accountable, therefore a computer must never make a management decision."
  • Neither Cisco's nor Extreme's launch materials address who's accountable when their agent makes a bad call at three in the morning — not a knock on either product specifically, but a sign the entire category is shipping capability well ahead of a governance answer

Deep Dive: This is the same fight this show has been tracking since GitLost — the GitHub Agentic Workflows credential-scope failure that leaked private repos to a public issue with zero attacker skill required. That story was about an agent having more standing access than the situation warranted. This one is different in kind: Cisco and Extreme aren't accidentally over-provisioning an agent, they're deliberately building the product category around agents holding real operational authority on production networks. NetClaw's maximalist tool-surface bet, ThousandEyes' minimalist counter-bet, and now two of the largest networking vendors on Earth publicly betting the network itself is the right layer for agentic control — the industry is converging on capability faster than it's converging on an answer to Willison's actual question.

Read Cisco's framing carefully: "the network is more powerful than the node" isn't just a tagline, it's an architectural claim that coordination and visibility matter more than any single model's intelligence — which, notably, is exactly the argument this show has made about automation and programmability for months. The interesting tension is that Cisco is right about where the value sits, and still hasn't said a word about who's on the hook when the Deep Network Model recommends a change that takes down a production fabric. That's not a product gap you fix with a better model. It's an organizational and contractual gap, and it's the same one Willison is pointing at from the opposite direction — a human, not the agent, remains the one who has to answer for it, just with a lot less visibility into why the call got made.

So What? If you're evaluating either platform — or building your own agentic NetOps layer on top of NetBox, Nautobot, or a custom stack — write the accountability answer down before you write the automation rule. Decide, in advance, who signs off when an agent-proposed change fails, and make sure the audit trail captures enough of the agent's reasoning that a human reviewing it after the fact isn't just staring at a diff with no context for why it happened.

SourcesCisco Newsroom, Cisco Blogs, Extreme Networks coverage, Simon Willison


3. netlab 26.07 Quietly Sets Up Distributed, Encrypted EVPN and VXLAN Labs

TL;DR: Ivan Pepelnjak's netlab shipped a multiserver plugin that spreads a single lab topology across more than one physical host, plus a WireGuard tunnel plugin for encrypting the links between them — the first real step toward labs that outgrow one machine and stay secure while doing it.

Key Points:

  • Multiserver plugin (contributor: @muddyblack) distributes containerlab devices across multiple physical hosts, removing the single-machine CPU and RAM ceiling that has always capped how realistic a netlab topology could get
  • WireGuard tunnel plugin (@jbemmel) adds encrypted point-to-point links — currently FRR-only — for securing traffic between distributed lab nodes
  • Also in this release: GRE tunnels across Cisco IOS, FRR, VyOS, and Junos; BGP and OSPF graceful restart on Arista EOS, BIRD, FortiOS, and FRR; and BIRD gains BGP confederations, VRFs, VXLAN, and EVPN MAC-VRFs — bringing BIRD to near feature parity with FRR for EVPN and VXLAN fabric testing
  • Breaking changes: the VirtualBox provider is dropped entirely (containerlab and KVM are now the only paths forward), and BIRD v3 becomes the default
  • Upgrade with pip3 install --upgrade networklab

Deep Dive: The multiserver plugin is the one that actually matters here. Realistic EVPN and VXLAN fabric topologies — multi-spine, multi-leaf, border-leaf, DCI, multiple route reflectors — routinely blow past what one host can run once you're using full-weight vendor images like cEOS or SR Linux rather than lightweight daemons. Multiserver is a direct fix for that ceiling, and it's what actually lets you test fabrics that look like what you'd run in production instead of a toy four-node topology. WireGuard is the complementary piece: on its own it's just an encrypted tunnel between two FRR nodes, but once you're spreading a fabric across multiple hosts — a home lab box plus a cloud VM, say — you need a secure transport for the inter-server underlay links, and that's exactly what it's for. It's FRR-only today, so it can't yet secure links terminating on a full vendor NOS node in the same fabric, which makes this the first step toward distributed, secure-by-default labs rather than a finished capability.

The quieter but genuinely useful bit is BIRD catching up to near-parity with FRR on EVPN and VXLAN. BIRD's resource footprint is much lower than FRR or any vendor image, so bigger fabrics now fit on the same hardware — which matters even more once multiserver is spreading things across machines that might not all be beefy.

So What? If you're building or maintaining EVPN/VXLAN lab environments for design validation or training, upgrade to 26.07 and specifically test the multiserver plugin against a topology you've been avoiding because it was too big for one box. If you're running Mikrotik in your labs, check the BGP template syntax before upgrading — it changed in this release.

SourcesipSpace.net


Networking
№ 02·Networking

Networking & Architecture

Plate IInetworking
Schematic leaf-spine fabric — explicit-path traffic flows across the spine plane, pods at the edges.

An AllReduce Algorithm Built for the Topology Google Actually Uses

TL;DR: A TU Berlin paper proposes Trivance, an AllReduce algorithm that completes in a third fewer communication steps than today's standard approaches by using both directions of a bidirectional ring at once — aimed squarely at torus-shaped AI training networks like Google's TPUv4, not the fat-tree/Clos topologies most enterprise fabrics use.

Key Points:

  • Claims 50% fewer steps than Swing and Recursive Doubling algorithms by tripling the effective communication distance covered per step
  • Packet-level simulation shows 5-30% improvement for AllReduce payloads up to 8 MiB, extending to 32 MiB in high-bandwidth settings and 128 MiB on 3D tori
  • Cuts congestion threefold versus Bruck's algorithm while staying bandwidth-optimal
  • Simulation-only — no real hardware or testbed results yet — but the paper was revised as recently as July 9th, so this is active, current work, not a stale preprint

So What? If you're anywhere near direct-connect or torus-topology AI training fabrics rather than fat-tree designs, this is worth tracking as it moves toward hardware validation — collective-communication overhead is one of the largest hidden taxes on GPU utilization at scale, and a genuine step-count reduction compounds fast at training-cluster size.

SourcesarXiv 2602.17254


Someone Reverse-Engineered Huawei's Proprietary AI Interconnect and Published the Spec

TL;DR: A single-author project called OpenURMA is a clean-room, open-source reimplementation of Huawei's Unified Bus — the proprietary interconnect behind Huawei's Ascend 950 AI accelerator fabrics and its answer to NVLink — built entirely from public documentation and reverse engineering, not access to Huawei's own hardware or code.

Key Points:

  • Reports roughly 500 nanoseconds of end-to-end latency for a 64-byte remote memory fetch — a claimed 4.37 times improvement over a matched RoCEv2 baseline at 2,186 nanoseconds — plus 2.8 times higher throughput
  • Fits in about 14% of an Alveo U50 FPGA's logic cells; evaluated across synthesizable RTL, SystemC, and gem5 full-system simulation
  • No ASIC or production deployment — this is FPGA and simulator evaluation from one academic author, so treat the performance numbers as a research claim, not an independently validated benchmark
  • Submitted May 27th; no public GitHub repository surfaced alongside the abstract

So What? The interesting part isn't whether the exact latency numbers hold up under scrutiny — it's that Unified Bus, like NVLink and UALink, is precisely the kind of proprietary interconnect standard that determines whether disaggregated AI fabrics stay vendor-locked. An open, clean-room implementation is a small but real crack in that wall, worth watching if you care about interconnect standards independent of any one vendor's roadmap.

SourcesarXiv 2605.28717

(This domain's automation crossover story — netlab 26.07 — is Top 3 story #3, above.)


Automation
Plate IIIautomation
Source-of-truth pipeline — intent → diff → apply → verify, idempotent on every revolution.

Deutsche Telekom Open-Sources Its Own Transport Automation Platform

TL;DR: Deutsche Telekom contributed StratoWeave, a declarative, vendor-agnostic platform for automating large-scale network configuration, to Linux Foundation Networking — a direct bet that the future of telco-grade automation is open rather than built on proprietary orchestration stacks.

Key Points:

  • Uses model-driven, declarative "transforms" — operators specify desired network end-states rather than step-by-step implementation, in the same spirit as intent-based config generation
  • Architecture supports layered service models, closed-loop automation, and reusable automation code for common operator services
  • "AI-enabled network operations" is stated as a direction, not a shipped capability — no published API spec, protocol list, or code walkthrough yet
  • Currently at LFN "Candidate" status — the earliest project tier, meaning this is an intent-to-build announcement more than usable software today

So What? Not something to adopt yet, but worth a bookmark if you work anywhere near service-provider automation — a Deutsche Telekom-scale operator choosing to open-source its internal platform rather than build further on Cisco NSO or a commercial orchestration stack is a real signal about where large telcos think GitOps-for-networking is heading. Revisit once there's an actual repository and API surface to evaluate.

SourcesLinux Foundation

(This domain's two biggest stories this cycle — the Cisco/Extreme agentic control-plane story and netlab 26.07 — are Top 3 stories #2 and #3, above.)


AI / ML
№ 04·AI / ML

AI & Machine Learning

Plate IVai / ml
Embedding space — clusters carry related concepts; the highlighted query vector pulls its nearest neighbors.

GPT-5.6 Splits the Frontier Three Ways — and Trails Claude on the Benchmark That Matters Most

TL;DR: OpenAI shipped GPT-5.6 as three named capability tiers — Sol, Terra, and Luna — replacing size suffixes with persistent tier names; on independent benchmarks the results are a genuine split decision against Claude, not the clean win OpenAI's own framing implies.

Key Points:

  • Pricing per million tokens: Sol at five dollars input/thirty dollars output, Terra at two-fifty/fifteen, Luna at one dollar/six dollars — a real price and efficiency play, not just a capability release
  • On SWE-Bench Pro — a benchmark OpenAI didn't design — Sol scored 64.6%, trailing Claude Mythos 5's 80.3% and Fable 5's 80% by roughly fifteen points
  • On the independently-run Artificial Analysis Coding Agent Index, Sol leads at 80 versus Fable 5's 77.2; on the Artificial Analysis Intelligence Index, Fable 5 leads 59.9 to Sol's 58.9 — genuinely mixed results depending on which third-party index you read
  • OSWorld 2.0: Sol hits 62.6% using 85% fewer tokens than Claude Opus 4.8 — a real efficiency claim worth separating from the raw-capability comparison
  • A self-reporting red flag: OpenAI's headlined "Agents' Last Exam" score of 53.6% doesn't match its own published table figure of 52.7%, and OpenAI doesn't disclose which reasoning configuration produced the higher number — flagged independently by both MarkTechPost and Simon Willison
Plate VSWE-Bench Pro — July 2026 frontier releases
SWE-Bench Pro — July 2026 frontier releases
vendormodelweightsSWE-bench
OpenAIGPT-5.6 Solclosed64.6%
AnthropicClaude Fable 5closed80%
AnthropicClaude Mythos 5closed80.3%

So What? Don't take either vendor's own framing at face value — OpenAI wins on price and some agentic-tool benchmarks, Anthropic still leads on SWE-Bench Pro and the Artificial Analysis Intelligence Index. The more durable signal for infrastructure planning is the tiering model itself (named capability tiers instead of version-number suffixes) and the new sandboxed "Programmatic Tool Calling" pattern — both are architecture decisions other vendors are likely to copy regardless of who's ahead this month.

SourcesOpenAI, MarkTechPost, Simon Willison

(This domain's governance angle — Simon Willison's "Directly Responsible Individuals" piece — is folded into Top 3 story #2, above.)


Datacenter
№ 05·Datacenter

Datacenter & Infrastructure

Plate VIdatacenter
Datacenter row — per-rack utilization at a glance. Cool colors are slack; warmer fills are pressure.

The Electrical Infrastructure Gap Has Real Numbers Behind It

TL;DR: A DataCenter Dynamics piece argues the actual AI infrastructure bottleneck isn't GPU allocation, it's a shortage of electrical-design expertise capable of specifying DC-ready switchgear, solid-state transformers, and high-ampacity cable architectures for GPU-dense racks — and backs it with real, checkable numbers rather than consultant framing.

Key Points:

  • Standard new AI facility racks now run 15-50kW; GPU-dense configurations run 100-250kW per rack
  • Most electrical contractors have deep experience with standard AC design; DC-capable design and commissioning talent — solid-state transformers, DC-ready breakers, new cable ampacity architectures — is genuinely scarce
  • The datacenter switchgear market is projected to hit $13.6 billion by 2031, a 16% compound annual growth rate
  • The framing: this is a scheduling risk, not a competitive edge — every operator is hitting the same talent shortage at the same time

So What? If your organization is anywhere near AI datacenter facility planning, treat electrical-design and commissioning lead time as a critical-path item alongside chip allocation and grid interconnection — it's the quieter constraint sitting underneath the power-availability headlines.

SourcesDataCenter Dynamics

(This domain's biggest story this cycle — DriveNets' scale-across supercluster — is Top 3 story #1, above.)


Science
Plate VIIscience
Field schematic — three-body stability under quasi-equal masses, drawn from the day's central result.

A "Boring" Liquid Just Cracked Like Glass, and Physicists Don't Fully Know Why

TL;DR: Drexel University researchers stretching a viscous, non-elastic hydrocarbon fluid during industry testing with ExxonMobil heard it audibly crack instead of just thinning and flowing — a brittle-fracture regime in a simple fluid that textbook fluid dynamics said shouldn't be possible without elasticity in the mix.

The Science: Using extensional rheology — pulling the fluid between plates at speeds up to 500 millimeters per second — the team measured crack propagation at 500 to 1,500 meters per second, versus roughly 0.07 meters per second in the elastic "complex fluids," like polymer melts, where fracture was previously known to occur. Critical fracture stress held consistent at 2 megapascals across samples; only the least-viscous samples failed to fracture at all. The leading explanation isn't elasticity — there essentially is none here — but cohesive molecular energy, reviving a theoretical prediction from fluid dynamicist Daniel Joseph in the 1990s about cavitation-driven fracture in simple fluids. Peer-reviewed in Physical Review Letters.

Why It's Interesting: This upends a basic assumption in fluid dynamics — that you need elasticity to get true fracture rather than just viscous necking and thinning. It's not just an academic curiosity: industrial processes that pump, spray, or stretch "simple" viscous fluids at high rates — coatings, adhesives, oil and gas extraction, hence Exxon's interest — now have to consider that these fluids can shatter rather than flow under the wrong conditions, which is a real process-design and failure-mode question, not a lab-only phenomenon.

SourcesQuanta Magazine


A 58-Year-Old Quantum Contradiction Just Got Resolved

TL;DR: Heidelberg University theorists unified two long-standing, seemingly contradictory pictures of how a single particle moves through a sea of fermions — the mobile "Fermi polaron" and the "Anderson orthogonality catastrophe," in which a heavy, frozen impurity was thought to destroy quasiparticle behavior entirely.

Key Points:

  • The resolution: even an extremely heavy, nominally motionless impurity still makes small residual movements as the surrounding fermion sea relaxes around it, and those tiny motions open an energy gap that lets quasiparticle behavior re-emerge rather than break down
  • Theoretical and analytical, not an experimental result — published in Physical Review Letters after first circulating as a preprint
  • Directly relevant to interpreting ultracold-atom, 2D-material, and correlated-electron experiments, anywhere a dopant or defect sits inside a sea of fermions

So What? A cleaner unified framework for when quasiparticles survive versus break down is a real, if quiet, win for anyone interpreting transport or spectroscopy measurements in quantum-simulation platforms.

SourcesScienceDaily, EurekAlert


Quick Takes
№ 07·Quick Takes

Quick Takes

  • Anthropic extended Claude Fable 5's free access on paid plans and kept Claude Code's elevated weekly rate limit through July 19th — the second extension in about a week, with metered usage pricing now slated to begin July 20th. Worth flagging as an operational note rather than industry gossip: this pipeline runs on Claude Code, so the cutover date is a real capacity and cost-planning item.
  • A new arXiv paper proposes verifying 6G RAN network intent using only aggregate, standardized telemetry counters — no deep packet inspection — to resolve the tension between multi-vendor O-RAN assurance and GDPR-style privacy rules; evaluated against real production management-plane data from four operator networks, which is more substance than most papers in this space bring.
  • Liquid cooling's M&A wave keeps building: Ecolab closed its roughly $4.75 billion all-cash acquisition of CoolIT Systems, and Eaton folded in Boyd Thermal as part of a reported $11 billion in acquisitions in a single quarter — real corporate acquisitions with disclosed prices, not the vague minority-stake or securitization deals this show usually has to squint at.
  • Friday's opposition-tracking story got a sharper edge: reporting since then puts specifics on the trend — more than 300 data center bills introduced across state legislatures in the first six weeks of 2026, fourteen states floating outright construction moratoriums, Norway capping new permits at 5 megawatts, and unanimous council votes against proposed projects in both Tucson and Indianapolis before any approval vote even happened.
  • Memory makers are riding the AI boom hard: SK Hynix and Micron revenue roughly tripled year over year, Samsung's roughly doubled — but The Register's read is that the industry's own boom-bust history says a reversal is coming, not a permanent plateau.

SourcesBleepingComputer, arXiv 2607.08809, Ecolab, Yahoo Finance, Brookings, The Register


Watch Today
№ 08·Watch Today

Watch Today

  • DriveNets' scale-across claims — watch for an independent third party or a primary-source confirmation of the 111.2 Tbps / 0.9 millisecond figures; we couldn't verify them directly this cycle.
  • StratoWeave's first commits — worth checking back once Deutsche Telekom's contribution has an actual repository and API surface to evaluate, not just a press release.
  • Whether Cisco or Extreme publishes an accountability model for their agentic platforms — neither launch addressed who signs off when the agent is wrong.
  • The Claude Code rate-limit cutover on July 20th — metered pricing begins then; worth checking this pipeline's own usage patterns against it before it lands.
  • Independent reproduction of GPT-5.6 Sol's benchmark claims, especially the "Agents' Last Exam" number that doesn't match OpenAI's own published table.

Automation
№ 09·Automation

Pipeline Stats

Plate VIIIautomation
Source-of-truth pipeline — intent → diff → apply → verify, idempotent on every revolution.
  • Domains researched: 6 (network architecture, network automation, AI/ML, security, science, datacenter)
  • RSS digest: thin cycle — 22 articles / 22 feeds, top relevance score 13.6 (netlab 26.07) — most of today's coverage came from targeted supplemental search given the thin digest
  • Items published: 9 primary items + 5 quick takes
  • Security: no significant architecture updates this cycle (checked zero-trust and microsegmentation sources directly; nothing dated within the last 24-48 hours cleared the bar)
  • Quality score average: 4.5 / 5
Subscribe

Get the briefing in your inbox.

One email per weekday morning. Same writing, same sources — no audio required.