Scale-up domain
Beginner
A team of AI chips joined by such fast links that they can work almost like one giant chip.
Novice
The set of accelerators joined by a dedicated high-bandwidth, low-latency fabric, usually inside one server or one rack. Within it, chips can read and write each other’s memory and run group operations quickly. Also called an NVLink domain, a pod or a node, depending on the vendor.
Expert
The largest set of accelerators that share a memory-semantic fabric with roughly an order of magnitude more bandwidth per chip than the scale-out NIC. Its size (8, 16, 64, 72 chips in current systems) bounds tensor and expert parallelism, which is why designers stretch it to rack scale over copper.
Explained in Scale-up fabrics (Systems).
See also: Scale-out network, Collective operation.