NVIDIA Ada Lovelace
NVIDIA L40S GPU Cloud & Dedicated Infrastructure
Versatile GPU for inference, rendering and generative media workloads.
The NVIDIA L40S is built on the Ada Lovelace architecture and is well suited to inference, image and video generation, and graphics-adjacent workloads. Its strong price-to-performance ratio for inference makes it a common choice for teams serving models in production without the cost of Hopper-class training GPUs.
L40S specifications
- GPU Memory
- 48GB GDDR6 with ECC
- Architecture
- NVIDIA Ada Lovelace
- Interconnect
- PCIe Gen4
- RT / Tensor Cores
- 3rd-gen RT cores, 4th-gen Tensor cores
- Form Factors
- PCIe, dual-slot
- Typical Server Configs
- 1x, 2x, 4x, 8x GPU nodes
Deployment options
- On-demand cloud GPU instances
- Reserved capacity
- Dedicated L40S servers
- GPU worker pools for generation queues
Best suited for
- LLM and diffusion model inference
- Image and video generation (Stable Diffusion, ComfyUI, FLUX workflows)
- Rendering and graphics-accelerated workloads
- Cost-efficient inference-only deployments
Less ideal for
- Large-scale distributed pretraining
- Workloads that require HBM-class memory bandwidth
Related solutions
L40S frequently asked questions
Is the L40S suitable for LLM inference?
Yes. The L40S is a popular choice for serving small to mid-sized LLMs and diffusion models where inference cost efficiency matters more than raw training throughput.
Can I run ComfyUI or Stable Diffusion workflows on L40S?
Yes. The L40S is well suited to image and video generation workloads, including ComfyUI-based pipelines. See our managed ComfyUI infrastructure for a production-ready deployment option.
How does L40S pricing typically compare to H100?
L40S capacity is generally more cost-efficient per GPU-hour than H100, reflecting its focus on inference and graphics workloads rather than large-scale training. Request current pricing for an accurate comparison based on your workload.