Introduction: Guide to Building AI Factories
About This Document
This NVIDIA white paper provides a framework for building single-tenant enterprise AI factories. It describes a starting reference architecture and an ecosystem approach that brings together accelerated computing, high-performance networking, storage systems, and AI software. The guide is intended for partners and enterprise teams responsible for infrastructure design, procurement, and operations.
Key Takeaways
- The document presents the AI factory as enterprise infrastructure that must be designed across compute, networking, storage, and software domains simultaneously.
- It proposes a single-tenant architecture designed around the needs of a single enterprise.
- NVIDIA positions its certified partner ecosystem as a means of integrating components and reducing risk in AI infrastructure projects.
- The target criteria for an AI factory are cost efficiency, scalability, and high performance.
- Key participants in the solution lifecycle include OEMs, ISVs, systems integrators, AI developers, MLOps, and IT and network administrators.
Practical Value for Data Center Owners
For data-center owners and project teams, the paper is useful as an initial framework for developing a dedicated AI-cluster concept: procurement and integration should be assessed not only in terms of GPU servers, but also the compatibility of certified servers, storage, networking, and AI software. The reference architecture can be used to define supplier requirements and allocate responsibilities among OEMs, the systems integrator, and IT and network teams.
Where It Applies
Topics
Source: NVIDIA · open page