Infrastructure · Infrastructure resource
Serverless Function Runtime
Infrastructure resourceInfrastructureInfrastructurearc:ServerlessFunctionRuntime
A managed execution platform that runs stateless functions on demand in response to events, scaling from zero to thousands of concurrent executions and billing per millisecond of compute.
Responsibility. Executes event-triggered agent functions without server management.
Also known as: Serverless platform, Function-as-a-Service
Variant of Agent Hosting Platform abstract
When to choose. Choose when traffic shows high variance or long idle periods, deployment velocity is critical, the team lacks infrastructure expertise, cost must scale with actual usage, or the system processes discrete events rather than continuous streams.
Relationships
hosts structural
is configured by structural
- Serverless Hosting Plan abstract Ch4.2
sends data to dynamic
alternative to variability
Design guidance
- SHOULD minimise cold-start latency by trimming dependencies (tree shaking), lazy-loading expensive resources and allocating more memory for CPU-intensive initialisation rather than enabling provisioned concurrency everywhere.
- SHOULD avoid VPC integration unless functions truly require private resources, because it adds substantial cold-start time.
- MUST externalise all state, since each invocation starts without memory of previous executions.
- SHOULD use persistent checkpoints or a durable orchestrator for workflows exceeding the platform's maximum execution time.
Quantitative guidance
As stated by the sources; verify before use.
- Maximum execution time: 15 minutes (Lambda), 10 minutes (Azure Functions); consumption-plan limit 5 minutes (Ch4.2).
- Execution environments stay warm 5-15 minutes; cold starts affect <1% of requests for moderately active functions (Ch4.2).
- VPC configuration adds 10+ s to cold starts; loading a 2 GB model on cold start adds 30+ s; optimisations cut cold starts from 5-10 s to <500 ms (Ch4.2).
- Billing example: 1M events x 200 ms x 1 GB = 200,000 GB-s vs a 730-hour server idle 90% of the time (Ch4.2).
- Worked example: 19,785,600 GB-s/month at $0.0000166667 = $329.76 + $0.90 requests = $330.66/month; 5-15% of requests see 500-800 ms cold starts; deployments complete in ~30 s (Ch4.2).
Classification
- Patterns
- Scale-to-zeroPay-per-executionEvent-driven executionWarm execution environment reuse
- Technologies
- AWS LambdaAzure FunctionsGoogle Cloud Functions
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Cost efficiencyMaintainability (ISO/IEC 25010)
- Risks mitigated
- Paying for idle capacity during low trafficInability to scale quickly during traffic spikesCapacity planning errors
Sources
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.