The 115-Petabit Lifeline: How Broadcom’s Jericho4 Flips the Script on Mega Datacenters
Ever wonder what powers the next leap in artificial intelligence? It might not be what you think. Giants like Meta, Google, and Microsoft are racing to build colossal, multi-gigawatt datacenters – minicities demanding their own power plants and facing community opposition and years of construction delays. But what if we could break free from this model? Enter Broadcom’s Jericho4 switch, unveiled this week. This isn’t just another networking chip; it’s a radical proposition: distributed AI training across datacenters up to 100 kilometers apart. Forget building monolithic power-hungry behemoths; Jericho4 claims it can weave existing or future smaller facilities into a virtually unified supercomputer, potentially reshaping the very infrastructure of the AI boom.
Beyond the Single Campus: The Crippling Cost of AI’s Power Hunger
The current AI gold rush runs on graphical processing units (GPUs), voracious consumers of electricity and space. Training massive models like GPT-4 or Gemini requires thousands, even tens of thousands, of these chips working in concert. Concentrating them demands:
- Massive Power Infrastructure: Datacenter campuses now routinely exceed 1 GW. Meta’s planned complex in Richland Parish, Louisiana, required a dedicated new power plant generating 2.2 GW – equivalent to powering nearly 1.7 million homes.
- Severe Real Estate and Grid Demands: Finding locations with sufficient land, permits, water (for cooling), and crucially, an electrical grid capable of delivering gigawatts without major upgrades is increasingly difficult and time-consuming. Delays can stall crucial AI initiatives.
- Economic and Environmental Pressure: The sheer scale drives enormous costs and significant environmental footprints, even when pursuing renewable energy. Cooling densely packed GPUs is increasingly inefficient.
This centralized model is hitting its limits. Amir Sheffer, from Broadcom, contextualizes Jericho4’s role: “If you’re running a training cluster and you want to grow beyond the capacity of a single building, we’re the only valid solution out there.”
The Hardware Engine: Demystifying Jericho4’s Raw Power
Jericho4 is specifically engineered as a Datacenter Interconnect (DCI) powerhouse. Forget radix or ultra-low latency for intra-datacenter traffic (where Tomahawk 6 shines). Jericho4 targets the pipe between datacenters:
- Hyper-Port Innovation: Each Jericho4 switch houses up to eight specialized “Hyper-Ports.” Each hyper-port aggregates four 800 Gigabit Ethernet (GbE) links, functioning logically as a single colossal 3.2 Terabit per second (Tb/s) port.
- Meet the Aggregated ASIC Power: The entire Jericho4 ASIC boasts an astonishing 51.2 Tb/s of aggregate bandwidth across all its switch and fabric ports. This immense internal capacity prevents bottlenecks as external traffic flows in.
- The Efficiency Edge: Hyper-Ports vs. ECMP: Simply bundling regular ports using techniques like Equal-Cost Multi-Pathing (ECMP) leads to inefficiency due to imperfect load balancing. Broadcom asserts its hyper-ports achieve 70% higher link utilization – meaning significantly more usable bandwidth per physical connection. This directly translates to cost savings and higher ROI on the expensive DCI fiber links.
- Petabit-Scale Connectivity: The system scales to configurations involving 36,000 hyper ports, enabling blinding 115.2 Petabits per second (Pb/s) of bandwidth between two connected datacenters. This scale is deliberately targeted at massive-scale AI:
- 115.2 Pb/s = 115,200 Tb/s.
- Connecting 144,000 GPUs (each needing an 800 Gbps = 0.8 Tb/s uplink) would require precisely 115,200 Tb/s of non-blocking bandwidth (144,000 * 0.8 Tb/s = 115,200 Tb/s).
Facing the Inevitable: Latency and the Speed of Light
Distributing workloads geographically solves the power and space crunch, but introduces a fundamental challenge: latency. For tightly coupled parallel tasks like AI training, where GPUs constantly exchange data, nanoseconds matter.
- Light’s Hard Limit: No switch, no matter how advanced, can circumvent physics. Light travels through fiber optic cables at roughly 200,000 km/sec. Over 100km, the fundamental round-trip propagation delay is almost 1 millisecond (ms) just for the photons.
- Additional Latency Layers: This propagation delay is just the start. Transceivers converting electrical signals to light (and vice versa) add overhead. Network protocols (like TCP/IP or RoCE) add more. Switch processing and buffering contribute further (“tail latency”). Total end-to-end latency can easily double or triple the theoretical minimum.
- Jericho4’s Mitigation Strategies:
- Deep HBM Buffers: High Bandwidth Memory buffers can temporarily store massive amounts of data packets during congestion spikes, preventing expensive packet loss which causes retransmissions and even higher latency.
- Advanced Congestion Control: Sophisticated algorithms proactively manage traffic flow on high-speed links to minimize queuing delays and buffer bloat.
However, these optimizations manage problems around the light-speed delay; they don’t eliminate it. This necessitates innovation at the application level.
The Software Counterpart: Low-Communication Training Paradigms
If we can’t make light faster, we must make AI training algorithms less dependent on constant, real-time communication across vast distances. Pioneering research offers pathways:
- Streaming DiLoCo: Google DeepMind’s “Streaming DiLoCo” (Distributed Low-Communication method) replaces constant, global synchronization with:
- Stratified Local Processing: Dividing GPUs into independent “work groups” within each datacenter. These groups train largely in isolation.
- Infrequent Global Synchronization: Instead of continuous chatter, groups strategically exchange parameter updates periodically (e.g., every few hours or after processing large data chunks). Think “batched knowledge sharing” instead of “constant conversation.”
- Quantization: Transmitting updates using lower-precision numerical formats (e.g., 8-bit floats instead of 16 or 32-bit) dramatically shrinks the volume of data sent over the 100km link when syncs do occur. Less data means lower bandwidth consumption and quicker transfers even at fixed speeds.
- Optimized Communication Scheduling: Intelligently staggering when different groups communicate or overlapping communication with computation minimizes the impact of sync latency on overall training time.
Table: DiLoCo Communication Reduction vs. Traditional Syncs
| Feature | Traditional Parameter Server/All-Reduce | Streaming DiLoCo Approach |
|---|---|---|
| Communication Type | Constant, fine-grained synchronization | Occasional, coarse-grained model syncs |
| Data Volume | High (full gradients, frequently) | Lower (quantized updates, less frequently) |
| Network Bandwidth | Demanding constant high throughput | Tolerant of bursts; peaks are manageable |
| Latency Sensitivity | Extremely high (delays stall all computation) | Significantly lower tolerance required |
| Suitability for DCI | Poor (requires near-zero latency, high BW) | Ideal (fits Jericho4’s high BW capabilities) |
Jericho4 provides the highway capable of carrying the data torrent required for large-scale training; innovations like DiLoCo leverage this highway efficiently despite the distance, making distributed training technically viable. It’s a hardware-software symbiosis.
Strategic Implications: Beyond Gigawatts – Agility and Sustainability
Jericho4 transcends being just a switch; it represents a potential inflection point:
- Unlocking Geographically Distributed or Existing Sites: Companies aren’t forced to build massive new campuses. They can leverage multiple smaller, potentially underutilized sites, sites with stranded power capacity, or partners in different locations.
- Accelerated Scalability: Building smaller datacenters is faster than erecting gigawatt-scale behemoths. Adding capacity becomes more modular – deploy another smaller pod and connect it via DCI. Time-to-market for new AI clusters shrinks.
- Environmental Opportunities: Spreading workloads allows tapping into diverse, local renewable energy sources (geothermal, hydro, wind, solar) beyond what a single massive grid might support, potentially lowering the net carbon footprint compared to a single coal/gas-powered mega-cluster.
- Operational Resilience & Cost Flexibility: Multiple locations offer intelligent workload placement for cost optimization (based on local energy prices) and enhanced resilience against regional outages or disasters. Work can potentially migrate.
- Market Creation: Enables specialized “AI Accelerator Pod” providers – companies building standardized, efficient GPU pods specifically designed to be interconnected into larger fabrics using technologies like Jericho4’s DCI, offering capacity-as-a-service.
Reality Check and Outlook: Not a Panacea, But a Pivot Point
Jericho4 isn’t magic. The physics of latency remains. Algorithmic shifts like DiLoCo are crucial, and managing extremely complex global systems software across geographically distant clusters presents non-trivial operational hurdles. Traditional DCI methods involving over-subscription (e.g., 4:1 or 8:1) to manage costs may persist, albeit at massively higher aggregate speeds. The solution may become hybrid: giant super-clusters where feasible and unavoidable, plus interconnected smaller sites for agility and overflow.
Conclusion
Broadcom’s Jericho4 isn’t merely pushing bandwidth boundaries; it’s offering a potential blueprint for the sustainable evolution of AI infrastructure. By enabling the practical, high-bandwidth interconnection of datacenters 100km apart, it confronts the monumental power and scaling constraints threatening the AI revolution’s pace. It mitigates—but doesn’t eliminate—the pervasive latency challenge, a problem increasingly tackled through novel, low-communication training algorithms championed by AI developers like Google DeepMind. The era of absolutely necessary multi-gigawatt datacenters isn’t over, but Jericho4 and the ecosystem it fosters provide a compelling alternative path: scattering the supercomputer. This means faster deployment, geographical flexibility, potential environmental gains, and new market opportunities – all fueled by a 115-petabit bridge. Will interconnected pods become the new normal, or will the massive campuses inevitably dominate? The network battle for AI’s physical future is intensifying. What’s your take: distributed cloud or concentrated powerhouses for AI’s next leap? Share your thoughts below!
Sources & Further Reading:
Original article at go.theregister.com


