NVIDIA Triton Inference Server is an open-source inference serving software that simplifies deploying trained machine learning models at scale. It supports multiple frameworks (TensorFlow, PyTorch, ONNX), GPU and CPU execution, model ensembles, and dynamic batching. It optimizes latency and throughput for production inference pipelines.
Use this profile to understand the building block briefly, place it in the model, and open related building blocks.
Technical building block: can be automated, integrated, or operated.
Concrete cog in the system that works inside larger relationships.
Triton Inference Server is an NVIDIA server that exposes trained models from different frameworks as scalable inference services.
Triton emerged at NVIDIA from the need to operate models from different machine-learning frameworks consistently and efficiently in production. First released as TensorRT Inference Server, it was later expanded into Triton Inference Server.
A model and its configuration live in a repository. Triton loads them, accepts requests through standard interfaces, and runs inference on CPUs or GPUs; metrics and batching support operations.
Models are exposed as network services.
Frameworks connect through appropriate backends.
Batching and parallelism increase throughput.
Triton connects trained models to production applications and standardizes heterogeneous model operations.
Where this building block is located in the topic model.
No structure path available.
Explore how this building block connects to concepts, methods, technologies, and tools.
These sources establish the term and its professional meaning.
All direct connections of the current building block in a compact text view.
This classification shows where the building block typically matters, how demanding it is, and what kind of impact it has in the model.
The level within the organization (enterprise, domain, team) at which the AssetBlock is applied.