Model Serving · Software component
Model Graph Exporter
Software componentModel ServingModelsarc:ModelGraphExporter
A model optimisation component that serialises a framework-specific model into a framework-neutral computation-graph format with declared dynamic input dimensions.
Responsibility. Decouples the model definition from its training framework so that quantizers and engine builders can consume it.
Also known as: ONNX export, Model export step
Relationships
receives data from dynamic
- Foundation LLM abstract Ch7.4 Ref7.01
produces lifecycle
Design guidance
- SHOULD declare batch and sequence dimensions as dynamic at export so the graph does not assume fixed input shapes.
- SHOULD run export on hosts with memory of roughly twice the model size, or split the model into components.
- MUST register or replace custom operators that lack interchange-format equivalents before export.
Quantitative guidance
As stated by the sources; verify before use.
- Export of a 7B model takes 2-5 minutes and needs memory of roughly 2x model size (Ch7.4).
Classification
- Patterns
- Framework-agnostic intermediate representationDynamic axes (batch, sequence)
- Technologies
- ONNXPyTorch torch.onnx.exportTensorFlow
- Quality attributes
- Flexibility (ISO/IEC 25010)Compatibility (ISO/IEC 25010)
- Risks mitigated
- Engine inputs fixed to a single shapeFramework lock-in of the optimisation toolchain
Sources
- Ch7.4: T. Nguyen, "TensorRT-LLM Fundamentals and Quantization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.4. ISBN: 9798244538229.
- Ref7.01: NVIDIA, "Best practices," NVIDIA TensorRT Documentation. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/deeplearning/tensorrt/latest/performance/best-practices.html