Sprechen Sie mit unserem Server-Gehäuse Ingenieure und Vertriebsteam
24.000 Abonnenten
Teilen Sie uns Ihre Anwendungsanforderungen, den Gehäusetyp, die Rackhöhe, das Motherboard, die GPU, die Laufwerksschächte, das Netzteil, die Kühlung, die E/A-Anschlüsse und die Bestellmenge mit. Unsere Ingenieure und unser Vertriebsteam werden Ihnen ein Standardmodell oder eine OEM/ODM-Konfiguration für Ihr Projekt empfehlen.
Ein KI-Gehäuse mit hoher Dichte sollte auf die gesamte Architektur aus GPU, CPU, Hauptplatine, Netzteil, Netzwerk, Speicher und Kühlung ausgelegt sein – und nicht nur auf die Rack-Einheiten.
NVIDIA dokumentiert aktuelle KI-Plattformen im Rack-Maßstab mit einem Strombedarf von deutlich über 100 kW und zeigt damit, wie weit die KI-Infrastruktur die herkömmlichen Annahmen zur Serverdichte mittlerweile hinter sich gelassen hat.
Der Abstand zur GPU ist nur die erste mechanische Überprüfung. Der Abstand zwischen den Grafikkarten, die Befestigung, der Biegeradius der Kabel, die Stromanschlüsse, der Luftströmungswiderstand und der Wartungszugang sind ebenso wichtig.
Die Kühlungsarchitektur sollte festgelegt werden, bevor die interne Anordnung im Gehäuse endgültig festgelegt wird.
Eine maximale GPU-Dichte ist nicht automatisch gleichbedeutend mit guter Technik. Ein etwas größeres Gehäuse kann eine bessere dauerhafte Rechenleistung bieten und die Wartung vor Ort erheblich vereinfachen.
Bei den Prototypentests sollten vor Beginn der Serienproduktion realistische Wärmebelastungen, das Verhalten der Lüfter, die Kabelführung, das Gewicht der GPU, die Konfiguration des Netzteils sowie die zu erwartenden Wartungsabläufe nachgestellt werden.
KI-Hardware hat die Bedeutung des Begriffs “hohe Dichte” verändert.”
Vor einigen Jahren begann ein Projekt zur Entwicklung eines Servergehäuses vielleicht mit den Abmessungen des Motherboards, den Laufwerksschächten, den PCIe-Steckplätzen, dem Format des Netzteils und der Rackhöhe. Diese Faktoren spielen nach wie vor eine Rolle. Ein KI-Computing-Projekt bringt jedoch eine weitere Schwierigkeitsebene mit sich: mehrere leistungsstarke Beschleuniger auf engstem Raum, sehr hoher Strombedarf, schwere Karten, dichte Kabelbündel, platzsparende Kühlkörper, schnelle Netzwerkverbindungen und Kühlkomponenten, die um jeden Millimeter Platz konkurrieren.
Das verändert die Arbeit völlig.
Die Frage lautet nicht mehr einfach nur:, “Können wir acht GPUs in dieses Gehäuse einbauen?”
Eine bessere Frage lautet:
“Können wir das gesamte System in dieses Gehäuse einbauen und es unter Dauerlast betreiben, ohne dass dabei thermische, elektrische, konstruktive oder wartungstechnische Probleme auftreten?”
Das ist die Designphilosophie, die hinter einem ernstzunehmenden Maßgeschneidertes AI-Servergehäuse.
Die Rack-Dichte im KI-Bereich hat eine neue Dimension erreicht
Aktuelle Rack-Scale-KI-Plattformen zeigen, warum der Gehäusebau mehr Beachtung verdient.
Die Referenzarchitektur GB300 NVL72 von NVIDIA umfasst 72 Blackwell-Ultra-GPUs und 36 Grace-CPUs. Laut NVIDIA ist das Rack flüssigkeitsgekühlt, verfügt über acht 33-kW-Stromversorgungsmodule und benötigt möglicherweise bis zu 142 kW für ein voll bestücktes Rack.
Das ist keine gewöhnliche Serverraumdichte.
In der separaten Dokumentation zu Mission Control von NVIDIA zur Energieverwaltung wird die Nennleistung im Rack mit 120 kW für den GB200 NVL72 und 135 kW für den GB300 NVL72, mit maximalen Knoten-Leistungsbudgets von etwa 6,7 kW bzw. 7,5 kW. Diese Zahlen werden im Zusammenhang mit der Leistungsbudgetierung angegeben und sollten daher nicht als austauschbar mit jeder maximalen Rack-Spezifikation betrachtet werden; sie veranschaulichen jedoch dieselbe technische Realität: Moderne KI-Rechenleistung kann bei sehr geringem Platzbedarf eine enorme Leistung erbringen.
Bezugspunkt
Veröffentlichte Abbildung
Was dies für die Konstruktion von Gehäusen bedeutet
NVIDIA GB300 NVL72
72 GPUs + 36 Grace-CPUs
Extrem dichte Anordnung von Rechenleistung, Netzwerkkomponenten, Stromversorgung und Kühlmittelleitungen
Anforderungen für das GB300 NVL72 im Vollrack-Betrieb
Bis zu 142 kW
Die Wärme- und Leistungsarchitektur darf nicht erst im Nachhinein berücksichtigt werden
Auslegungsleistung des GB300-Racks im NVIDIA-PRS-Beispiel
135 kW
Die Stromverteilung und die Reserve müssen auf Systemebene berücksichtigt werden
GB300 – Maximale Knotenleistung im PRS-Beispiel
7,5 kW
Einzelne Recheneinheiten können eine erhebliche thermische Belastung aufnehmen
IEA: Beschleunigtes Wachstum des Stromverbrauchs durch Server
~30% jährlich im Basisszenario
Beschleunigtes Rechnen mit hoher Dichte dürfte sich zunehmend durchsetzen
Der allgemeine Trend geht in dieselbe Richtung. Die Analyse „Energy and AI“ der Internationalen Energieagentur prognostiziert, dass der weltweite Stromverbrauch von Rechenzentren etwa 945 TWh bis 2030 im Basisszenario. Der Stromverbrauch durch beschleunigte Server, der hauptsächlich durch den Einsatz von KI getrieben wird, wird voraussichtlich um etwa 30% pro Jahr, im Vergleich zu etwa 9% bei herkömmlichen Servern.
Für Gehäusedesigner und OEM-Einkäufer ist das von Bedeutung.
Leistungsstärkere Hardware bedeutet, dass es immer mehr Projekte gibt, bei denen Leistungsdichte, GPU-Auslastung, Kühlleistung, Kabelmanagement und Wartungszugang zu mechanischen Anforderungen erster Ordnung werden.
Beginnen Sie mit der Belastung, nicht mit dem Blech
Eine der schnellsten Möglichkeiten, ein Projekt für ein maßgeschneidertes Gehäuse zum Scheitern zu bringen, besteht darin, mit der CAD-Arbeit zu beginnen, bevor die Rechnerarchitektur feststeht.
Tu das nicht.
Before chassis geometry is frozen, the engineering team should know the intended workload and the major components expected inside the system.
At minimum, lock down:
GPU model and quantity
GPU dimensions and slot width
GPU power connector location
Motherboard model and drawing
CPU platform and cooler geometry
DIMM height and keep-out zones
PCIe risers or switching architecture
NVMe and storage requirements
Network adapters and interconnects
PSU type, quantity, and redundancy
Maximum expected system power
Fan size, thickness, pressure requirement, and control method
Air or liquid cooling strategy
Radiator, manifold, CDU, or hose requirements where applicable
Front and rear I/O
Rack depth
Rail arrangement
Maintenance direction
Target production quantity
Eine AI Server Enclosure should be engineered around this combined package. Changing one part later can trigger a chain reaction through the rest of the design.
Swap the GPU?
The connector location may move.
Change the PSU?
The cable bundle changes.
Add another network card?
Now the riser geometry changes.
Increase the fan thickness?
Suddenly the GPU power cable has nowhere comfortable to bend.
Millimeters pile up quickly.
GPU Fit Is More Than Length × Height × Width
This sounds obvious, but it causes expensive mistakes.
A GPU fitting inside the chassis does not mean the GPU works inside the chassis.
When evaluating a GPU-Server-Gehäuse, we need to look beyond nominal card dimensions and check the real installation envelope.
That includes the card body, heat sink, power connector, connector bend radius, adjacent card clearance, motherboard slot position, retention hardware, riser location, fan wall, cabling, and the direction in which a technician needs to remove the card.
Heavy accelerators create another issue: gravity.
Long cards can apply substantial leverage to the PCIe connection and rear mounting area. Shipping vibration makes that worse. A GPU that survives stationary operation on a lab bench may behave very differently after international freight, rack installation, repeated service cycles, or vibration from high-speed fans.
For dense systems, GPU retention should be part of the chassis structure—not an accessory someone remembers two days before shipment.
Cooling Architecture Should Be Frozen Early
Heat does not care how good the CAD drawing looks.
A dense server can be mechanically perfect and thermally awful.
For air-cooled systems, the chassis needs a deliberate pressure path. Air should enter where the hardware expects it, pass through the high-resistance components, and leave the enclosure without repeatedly circulating through hot zones.
That sounds simple.
Das ist es nicht.
Cables sit in the way. Drive cages sit in the way. PSU housings sit in the way. Tall DIMMs, risers, structural beams, radiator brackets, connectors, filters, backplanes, and even badly positioned sheet-metal flanges can steal pressure from the airflow path.
A High-Density AI Rack therefore cannot be designed by counting fan positions alone.
“Six fans” tells me almost nothing.
I want to know fan dimensions, fan curves, static pressure, restriction, inlet temperature, exhaust path, control strategy, component impedance, and what happens when one fan stops.
Those are useful questions.
Luftkühlung vs. Flüssigkeitskühlung
Design Factor
Luftkühlung
Flüssigkeitskühlung
Mechanische Komplexität
Unter
Höher
Chassis airflow dependency
Sehr hoch
Reduced for liquid-cooled components
Facility integration
Usually simpler
May require CDU, manifold, hoses, facility water connection
Leak-management requirement
No coolant loop
Ja
Service procedures
Familiar to most technicians
Requires coolant-aware service procedures
Suitability as heat density rises
Increasingly difficult
Often more practical
Internal space competition
Fans and ducts
Cold plates, tubing, manifolds, connectors
Failure planning
Fan redundancy and airflow loss
Pump, CDU, leak, flow, and facility-loop scenarios
The technician needs enough room to disconnect and reconnect a loop without removing unrelated hardware or soaking electronics.
NVIDIA’s GB300 NVL72 architecture itself uses rack-level and tray-level leakage detection alongside liquid cooling, which is a useful reminder that coolant management is part of the system architecture, not simply a cold plate bolted onto a GPU.
Power Density Changes the Mechanical Layout
Power distribution is often treated as an electrical topic.
Inside a dense enclosure, it is also very much a mechanical topic.
High-current power means larger connectors, more cables, tighter bend limitations, additional busbar considerations, higher connector temperatures, and less freedom to route wiring through whatever space remains after the “important” components are installed.
The power architecture needs physical territory.
Reserve it.
A chassis designed around multiple accelerators may need redundant power supplies, high-current distribution, accessible connectors, safe cable separation, airflow around the PSU zone, and room to replace a failed module without pulling the server apart.
Dies ist der Ort, an dem ein Servergehäuse mit hoher Dichte can go wrong even when the component list looks perfectly compatible on paper.
Everything fits.
Nothing is serviceable.
The 30–40 kW Rack Story That Changed How I Think About Density
Recently, while reviewing industry discussions, I came across an older r/sysadmin thread about high-density rack cooling that stuck with me.
The original poster had several NVIDIA DGX systems plus storage equipment and was exploring how to fit roughly 30–40 kW into one rack. One participant described operating cabinets at up to about 36 kW and said the margin during cooling failure could be brutally small—around 20 seconds before thermal shutdown for many systems in that specific deployment. Other engineers in the thread repeatedly raised the same concerns: cooling availability, facility power, emergency shutdown, plumbing, maintenance expertise, and the risks created by concentrating so much heat in one cabinet.
That discussion caught my attention because buyers sometimes start a custom AI project with one seductive question:
How much compute can we squeeze into this space?
Wrong first question.
The better one is:
How much compute can this enclosure support continuously, safely, and serviceably under the real facility conditions?
A chassis can look fantastic on a CAD screen. Every component fits. Every rack unit is occupied.
Then the workload starts.
Fans ramp.
Cable temperatures rise.
GPU exhaust collides with another heat source.
One component throttles.
A technician tries to remove a card and discovers that three cables and a manifold are blocking it.
That is not high-density engineering.
That is high-density packaging.
There is a difference.
Here Is the Part Many Buyers Don’t Like Hearing
Maximum component density is often a bad design target for a custom AI enclosure.
Yes, I said it.
“Eight GPUs in 4U” looks great in a comparison table.
“Maximum compute per rack unit” looks great in a procurement presentation.
But one more GPU is worthless if adding it creates poor airflow, thermal throttling, inaccessible connectors, overloaded power routing, difficult field repairs, or zero margin when room conditions drift away from the perfect laboratory assumption.
The winning enclosure is not automatically the smallest one.
For many projects, I would rather build a slightly larger chassis with predictable airflow, solid GPU support, clean power distribution, logical cable routes, accessible fans, and enough room for technicians to work than win a density contest that makes the finished system miserable to operate.
The best high-density enclosure is the one that can sustain the workload.
Not the one that wins the Tetris game.
Serviceability Has to Be Designed In
Ask a simple question during CAD review:
What fails first, and how do we replace it?
Then physically simulate the answer.
Can the technician replace a fan without removing GPUs?
Can the PSU come out from the rear?
Can a network adapter be accessed without disturbing coolant hoses?
Can a GPU be removed vertically or horizontally without disconnecting unrelated cables?
Can the front panel be serviced without pulling the complete server?
Can a leaking fitting be reached quickly?
Can a failed boot drive be swapped without touching the compute section?
These questions are not glamorous.
They save money.
For a system integrator building 50, 100, or 500 machines, a few extra minutes of service time per unit becomes real operational cost. A poor access sequence also raises the chance that technicians damage nearby cables, connectors, cards, or tubing while fixing something unrelated.
That is why a Custom Rackmount Enclosure should be designed not only for assembly but also for disassembly.
Production engineers build it once.
Your customer may service it for years.
Structural Design Matters More as GPUs Get Bigger
AI servers are heavy.
Sometimes extremely heavy.
The chassis has to manage that load during assembly, transport, rack insertion, normal operation, and maintenance. Long GPU cards, multiple PSUs, large motherboards, copper heat sinks, radiators, manifolds, and dense storage can all shift the center of mass.
Prüfen:
Sheet-metal thickness
Bend geometry
Local reinforcement
GPU-Halterungen
Unterstützung für Netzteile
Rail mounting locations
Handle loads
Rack insertion forces
Chassis sag
Torsional stiffness
Shipping orientation
Drop and vibration risks
A bracket may look insignificant in isolation.
Multiply that tiny deflection across a long heavy card, add shipping vibration, then repeat it across a production batch.
Now it matters.
Prototype Testing Should Try to Break Your Assumptions
A prototype is not just a metal sample for checking whether the screw holes line up.
Use it aggressively.
Load the actual motherboard.
Install the actual GPUs.
Use production-equivalent power cables.
Install the same risers, drives, network adapters, fans, radiators, tubing, PSUs, rails, and front-panel assemblies expected in the finished server.
Then test the ugly conditions.
Mechanical Validation
Check component installation, card retention, cable routing, connector access, chassis stiffness, rail engagement, lid fit, and service sequence.
Thermal Validation
Run representative workloads. Measure inlet and exhaust temperatures, component temperatures, fan behavior, hot zones, and performance under expected ambient conditions.
Power Validation
Confirm PSU loading, redundancy behavior, cable temperatures, connector temperatures, power distribution, and startup behavior.
Failure Validation
What happens when one fan stops?
What happens when a coolant pump or facility loop has a problem?
What happens when one PSU fails?
What happens if a filter becomes partially obstructed?
What happens when ambient temperature increases?
Shipping Validation
A machine that passes a thermal test but arrives with GPUs shifted inside the chassis is still a failed design.
Test packaging, retention, vibration resistance, brackets, fasteners, and heavy-component support before volume shipment.
What Should Be Included in the RFQ?
A weak RFQ produces assumptions.
Assumptions produce revisions.
Revisions cost time.
For a custom AI chassis project, send the supplier enough information to evaluate the complete system architecture rather than quote an empty metal shell.
A useful RFQ package should include:
Angaben zur Angebotsanfrage
Information to Provide
Formfaktor
Target U height, width, depth
Hauptplatine
Model, dimensions, drawings, mounting locations
GPU
Exact model, quantity, dimensions, slot width, power
The more complete the input, the earlier mechanical conflicts can be found.
Finding a 12 mm interference in CAD is cheap.
Finding it after tooling, fabrication, assembly, and international shipping is not.
Build Around Sustained Compute, Not a Density Number
High-density AI systems force mechanical, thermal, electrical, and facility engineering to meet in one box.
That is why a Maßgeschneidertes AI-Servergehäuse should never be specified by rack height and GPU count alone.
Start with the complete bill of materials.
Model the heat.
Reserve space for power.
Support the GPUs.
Plan the cable paths.
Decide the cooling architecture early.
Design access around the parts most likely to need service.
Then prototype the actual system—not an empty enclosure.
AI compute density will keep rising. The IEA’s projections for accelerated-server electricity use and today’s 100 kW-plus rack architectures already point in that direction.
The manufacturers and system integrators that handle that density well will not be the ones who simply make smaller boxes.
They will be the ones who understand what has to happen inside those boxes when the workload hits 100%.
FAQs
Was ist ein maßgeschneidertes KI-Servergehäuse?
Short answer: A custom AI server chassis is an enclosure engineered around a specific combination of GPUs, CPUs, motherboard, storage, networking, power, cooling, and rack requirements.
Unlike a generic server case, its internal layout can be optimized for accelerator spacing, structural support, airflow, liquid cooling, cable routing, redundant power, and maintenance access.
What makes a server chassis “high density”?
Short answer: A high-density server chassis places a large amount of compute, storage, or accelerator hardware into a limited rack space while maintaining adequate power delivery, cooling, structural support, and service access.
Density should be evaluated by usable sustained performance, not component count alone.
Ist 4U für GPU-Server immer besser als 5U oder 6U?
Short answer: No. A smaller chassis can improve rack density, but 5U or 6U may provide better cooling capacity, GPU spacing, cable routing, expansion, and maintenance access.
The correct form factor depends on GPU type, power, cooling strategy, motherboard layout, and operational requirements.
Wann sollte ein KI-Server eine Flüssigkeitskühlung verwenden?
Short answer: Liquid cooling becomes attractive when heat density, accelerator power, rack density, noise limits, or airflow restrictions make conventional air cooling difficult to manage efficiently.
The decision must also consider facility water, CDU requirements, manifolds, leak management, service procedures, and redundancy.
Welche Informationen benötigt ein Gehäusehersteller, bevor er ein KI-Gehäuse entwirft?
Short answer: Provide the GPU, motherboard, CPU, PSU, storage, networking, PCIe, cooling, rack dimensions, I/O, rail requirements, production volume, and target market.
Exact component drawings are far more useful than generic statements such as “8-GPU AI server.”
Warum ist die GPU-Bindung in einem Server mit hoher Dichte wichtig?
Short answer: Heavy GPUs can place mechanical stress on PCIe slots, brackets, and chassis structures during operation, shipping, and maintenance.
Proper retention controls card movement, distributes mechanical load, and reduces the risk of connector or board damage.
Sollte die maximale Anzahl an GPUs das Hauptziel bei der Entwicklung sein?
Short answer: Usually not. The better target is the highest practical compute density that can maintain thermal performance, reliable power delivery, structural stability, and acceptable service access.
One extra GPU is not valuable if the resulting system throttles or becomes difficult to maintain.
Mark Lee – Gründer und Spezialist für Servergehäuse (OEM/ODM)
Mark Lee ist der Gründer von ISTONECASE und verfügt über 20 Jahre Erfahrung in der Servergehäuse-Branche. Er ist spezialisiert auf OEM-/ODM-Lösungen für GPU- und KI-Gehäuse, Rackmount-, Industrie-, Wandmontage-, NAS-, Mini-ITX- und Multi-Node-Gehäuse. Mit seinem Fachwissen unterstützt er maßgeschneiderte Hardwareprojekte für Rechenzentren, KI-Computing, Unternehmensspeicher, Edge-Computing, Netzwerke und industrielle Anwendungen.