Supercomputing News logoSupercomputing News logoBeta
AIHPCQuantumEmerging
Subscribe
Supercomputing News logoSupercomputing News logo
Pillars
AI—HPC—Quantum—Emerging—
Theme
Subscribe
Supercomputing News logoSupercomputing News logo

Trusted reporting on AI, HPC, Quantum, and the technologies shaping the future of computing. Cryptographically signed. Agent-accessible.

Pillars

  • Artificial Intelligence
  • High-Performance Computing
  • Quantum Computing
  • Emerging Technology

Entities

  • Organizations
  • Products
  • People
  • Places

Publication

  • About
  • Contributors
  • Topics
  • Contact
  • For Agents

Weekly Update

Keep track of the biggest stories in supercomputing, every Thursday.

Subscribe for free today
© 2026 Supercomputing News
Privacy PolicyTerms of Use
Emerging TechnologyEmergingNews

Panmnesia and Meta Researchers Outline a Cross-Rack CXL Fabric for AI Data Centers

Panmnesia says the Review's design spans up to 960 accelerators in one coherent domain. Public sources say it describes component silicon, not a deployed system.

Illustration of a chip die floorplan on a dark background in which the functional blocks are rows of server racks, joined by a regular 4-by-4 grid of switch nodes and dotted traces. One cluster at lower left is drawn in solid white with lime green traces; the rest of the floorplan is a pale grey outline.
A cross-rack CXL fabric drawn as a chip floorplan: CPU trays, accelerator pods and memory pods laid out as blocks, joined by a fixed-hop switch grid. The solid cluster stands for the switch and controller silicon Panmnesia says is validated; the outline is the proposed data-center-scale architecture.AI-generated / SCN
SCN Staff
The Squad
Published
Sep 10, 2026
Add Supercomputing News as a preferred source on Google
Reading0%
Listen to this article12 min
Loading audio…
0:00/ 12:21PausedMuted
Played in full
Audio unavailable
Speed
1×
Download audioMP3 · 11.3 MB
0:00

Panmnesia, a fabless semiconductor company in Daejeon, South Korea, said on September 8 that it and Meta have proposed an AI data center architecture in which "an entire datacenter operates like a single chip." The proposal is a Review article in Nature Reviews Electrical Engineering, "One-chip-like datacenter design enabled by CXL-based scale-up fabrics," published online on August 10 and carried in the journal's September issue. Twelve authors are listed: ten from Panmnesia and two from a Meta group identified as "Meta Infra." NREE Reviews commissions the work, and Panmnesia says it is the first semiconductor startup to lead one. KAIST announced the work on September 9 as research by one of its professors and his faculty-founded startup, conducted with Meta researchers.

The architecture puts CPUs, accelerators, and memory into one Compute Express Link (CXL) domain and extends that domain beyond the rack. CXL is an open standard built on the PCIe physical layer. It lets a processor and an attached device share memory while hardware keeps their caches coherent. So far, its production use is mostly memory expansion and pooling inside a server, as SCN's June coverage of CXL pooling moving from demo to deployment described. Between racks, AI clusters run on network fabrics such as Ethernet or InfiniBand, and the time a cross-rack request takes varies from one request to the next. The Review's abstract calls this "latency spread" and says it "increasingly limits system-wide execution" as models scale across thousands of accelerators. Panmnesia told Blocks & Files that measurements of such environments "show heavy-tailed latency distributions, with 99th-percentile round-trip latency roughly five times the median."

What the Review proposes

The key points name three hardware structures a "truly chip-like datacenter" requires: high-fan-out non-blocking switches, link acceleration units, and fabric controllers. A non-blocking fabric is designed to carry simultaneous traffic between distinct input and output pairs within its capacity; contention for the same destination and flow-control delays can still add time. High fan-out means many ports per switch, so a path crosses fewer of them. The release credits the three elements together with bounded latency variability. In Panmnesia's description to Blocks & Files, the switch connects many devices at once so path lengths stay similar whatever the source; the LAU moves repetitive protocol processing at each connection point onto a dedicated hardware pipeline; and the fabric controller applies one ordering policy to requests across the whole system. The company told the publication that the fabric controller is a combined CXL/PCIe controller, that it and the LAU have completed silicon validation, and that the fabric switch has been fabricated, with pre-release silicon now being supplied.

Weekly Update

The biggest stories in supercomputing, once a week.

AI, HPC, quantum, and emerging tech. Reported, not aggregated.

Free · no account · unsubscribe anytime

Panmnesia's product pages, reviewed by SCN on September 10, list a link acceleration unit (LAU) IP and a Link Controller IP, both described as silicon IP for PCIe 7.0 and CXL 4.0. The Link Controller page lists 128 GT/s per lane, up to 16 lanes, and a round-trip latency below 100 nanoseconds. The LAU page describes hardware-managed coherence for CPUs, accelerators, and memory devices through back-invalidate snooping, snoop filtering, and a device coherence engine.

The glossary attached to the article defines directory-based coherence, snoop filtering, credit-based flow control, floorplanning and network-on-chip, the vocabulary of keeping caches consistent and traffic orderly inside a single processor. The second key point is that CXL "extends shared address space and hardware-managed coherence beyond CPUs, but its server-centric heritage prevents it from fully achieving chip-level execution at scale." The proposal is to rebuild those on-chip mechanisms as dedicated fabric hardware.

Deployment follows the same analogy. Resources are grouped by function into trays, trays into pods, and pods across a fabric, which the key points describe as "enabling fixed-hop, uniform-latency communication across racks." The authors propose the fixed-hop layout alongside the hardware mechanisms as a way to reduce latency variability; hop count controls one source of variation, and the hardware controls the others.

The numbers, and where they come from

Panmnesia's release states the effects against "a conventional rack-scale configuration" of one CPU coupled to two accelerators. By the company's account of the Review, a single CPU coordinates 16 accelerators rather than two; the coherence domain that operates as one unit grows to as many as 960 accelerators, "roughly 13 times the reference platform"; cross-rack round-trip latency falls "from the microsecond range to several hundred nanoseconds"; and the unit replaced after a failure narrows from a whole server to a single device. KAIST's announcement gives the same scale and latency figures and describes the 13x as a comparison with a current NVLink-based rack. Blocks & Files, describing the comparison, names NVIDIA's GB200 NVL72, in which one CPU is coupled to two accelerators over NVLink-C2C, as the reference configuration.

The Review's acknowledgments say Panmnesia supplied controller and LAU silicon and thank a colleague "for assistance with system evaluation." The company says it has implemented the core components in silicon and validated them. The release's own verb for the figures is "to gauge the effect of the architecture," and TechNode Global, reporting the announcement, described the figures as "architectural comparisons presented by the authors."

The lead author's laboratory at KAIST draws the same line in its own words. Its September 9 news entry calls the work "an architectural direction supported by core-component silicon validation, not an announcement of a fully deployed datacenter-scale system."

The lab's published hardware-scale figure comes from different work. In June, Panmnesia presented an ISCA 2026 industry-session paper on a CXL controller and port-based-routing switch, and the lab's June 29 entry reports "stable performance demonstrated at as many as 64 nodes."

On optics, the release says the Review "also sets out how optical links (CXL-over-optics) can extend the reach of a CXL fabric" and that "details are given in the Appendix." The abstract puts optics in the future tense, and the final key point says the design improves latency stability and coherence-domain size "within electrical link limits, yet inherently scalability is constrained," with "electrical-optical integration" listed under future progress.

Meta's part

Two Meta-affiliated researchers co-authored the Review. Kayvon Shakeri and Han Wang are listed under "Meta Infra, Meta, Menlo Park, CA, USA," and the contributions statement credits them, along with four Panmnesia authors, with having "contributed to the discussion of content." Lead author Myoungsoo Jung, Panmnesia's CEO and a KAIST professor, is credited with writing the article; four other Panmnesia authors share writing credit. The competing-interests declaration names the ten Panmnesia employees. The release does not quote Meta.

What Meta has published about its own infrastructure is narrower than the Review. Vistara, presented at ISCA 2026, describes production CXL memory expansion using DDR4 recovered from decommissioned servers and attached as CXL Type-3 memory to newer machines. In its evaluated workloads, disaggregated inference used up to 25 percent fewer servers, and one distributed-cache workload's average query-processing time fell 29 percent. That is memory expansion inside a server. Between data-center fabrics, Meta's February engineering post on backend aggregation describes "a centralized Ethernet-based super spine network layer," the design its engineers also discussed in SCN's coverage of scale-across networking for multi-building clusters. Co-authoring the Review is a research contribution, not evidence of what Meta intends to build.

Where it sits against the other cross-rack options

The first hop beyond the rack is a segment every incumbent is also working on. At Hot Chips 2026, NVIDIA's frame was a three-way split, as SCN reported: "NVLink for scale-up inside a rack, Spectrum-X Ethernet for scale-out across racks, and BlueField-4 for scale-in infrastructure offload." The Ultra Ethernet Consortium's 1.0 specification, released June 11, 2025, standardizes an Ethernet-based stack for scale-out. Microsoft and OpenAI's MRC paper describes a 75,000-GPU pretraining job over Ethernet, as SCN reported in May; MRC is its own transport work, separate from the UEC specification. The other open scale-up standard, UALink 1.0, released in April 2025, specifies "200G per lane scale-up connection for up to 1,024 accelerators within an AI computing pod," a pod size in the same range as the Review's 960.

The proposal extends CXL-based memory sharing and hardware coherence across racks. NVL72's 72-GPU NVLink domain and UALink's 1,024-accelerator pod (read, write, and atomic transactions over a load/store protocol, per the specification) are counts, and counts alone do not establish equivalent coherence semantics or workload performance. KAIST's announcement states the distinction the authors draw: NVLink and UALink, it says, focus mainly on connections inside the rack, while this architecture links CPUs and memory as well as accelerators over CXL and extends the range to the whole data center.

The CXL 4.0 white paper (November 2025) describes 128 GT/s on the PCIe 7.0 physical layer, native x2 links "to support increased fan-out" and "support for up to four retimers for increased channel reach beyond CXL 3.0." The Review places CXL-over-optics in future work. If an implementation used co-packaged optics, the constraints SCN reported on September 4, from trade-press accounts of TSMC's August 31 remarks on lasers, fiber, connectors, and test, would apply.

There is also a software question. The abstract frames the problem as "network-based distributed systems relying on software to coordinate separate devices," and the answer is to move coordination into hardware coherence. The next question is what runs on top of 960 hardware-coherent devices, from the scheduler down to the collective library. SCN's coverage of why CUDA's advantage has become a composability problem reminds us that hardware gains must be matched in software before they show up as throughput.

The proposal arrives as NVIDIA signs accelerator vendors onto NVLink Fusion. In the same week, d-Matrix, which bought GigaIO's PCIe memory-fabric business in April, said it would connect its Raptor inference chip to NVIDIA's rack architecture through NVLink Fusion, with initial availability expected in the fourth quarter of 2027. Panmnesia, a KAIST faculty startup, is proposing that an open standard become the cross-rack coherent fabric of the AI supercomputer. The authors do not present that as an either/or: the lab's research page lists "how CXL can work with accelerator-centric links such as UALink and NVLink" among its subjects.

Panmnesia's announced milestones

Panmnesia's releases describe the switch at the center of this fabric. In November 2025, the company announced sample availability of a PCIe 6.0/CXL 3.2 fabric switch, with samples going to "early access partners." In March 2026, it announced a partnership with SK Telecom on a CXL-based AI data center architecture that the two companies "plan to validate" by running real AI models "by the end of this year," with proof-of-concept deployments to follow. In June, the ISCA paper and the lab's 64-node figure appeared; the same entry says pre-release fusion switch chips are available. On September 9, the Review release described the core components as implemented in silicon, validated, and "now preparing ... for commercial supply."

SK Telecom is described as a development partner. The dated items are the company's own: the SK Telecom validation, targeted for the end of 2026, and the commercial supply that the September release says it is preparing.

AI InfrastructureSemiconductor ManufacturingOptical InterconnectsNetworkingSemiconductor Fabrication
AI disclosure
This article was prepared with AI assistance for research and drafting under human direction and editorial control, per SCN house style.
About the contributor
SCN Staff
The Squad

The SCN Staff is a small AI editorial squad working under human direction. Each agent owns one job.

Scout does the research. It runs down primary sources and checks what's already been published, on SCN and everywhere else, before a story gets written. If a claim can't be traced back to a real document, Scout flags it.

Forge writes. It takes what Scout found and turns it into a draft, argument and sentences and all. Every SCN piece starts here, then gets sharpened.

Cipher handles search: the titles, descriptions, and keyphrase work that decides whether a good article ever gets found. Least glamorous job on the squad. Also one that matters more than it looks.

Pixel makes the visuals. Images, charts, the occasional diagram, all built to SCN's brand instead of pulled from a stock library. When something's easier to see than to read, it goes to Pixel.

Editorial judgment and the final call stay with the humans. So does the fact-checking.

Related reading
AI · AnalysisApple's Mac Shortage Signals Memory Supply Chain Has Reorganized Around Data Center AIHPC · AnalysisORNL's Next-Generation Data Center Institute: National Lab Expertise Meets the AI BuildoutHPC · AnalysisSlingshot Held Performance Under AI Traffic Patterns That Collapsed InfiniBand by 5x on Production Exascale