

Best NVIDIA H100 cloud providers for AI inference in 2026
Renting NVIDIA H100 capacity is only one part of running production AI training and inference. This article compares application platforms, specialist GPU infrastructure, hyperscalers, and serverless execution across pricing, topology, infrastructure control, scaling, and production workload fit.
- Northflank suits engineers, startups, and enterprises deploying H100 workloads alongside APIs, workers, jobs, databases, CI/CD, autoscaling, and observability on managed infrastructure or through self-serve BYOC.
- Runpod suits teams that want flexible access to H100 PCIe, SXM, and NVL GPUs through Pods, Serverless, and Clusters.
- Lambda suits teams that want dedicated H100 infrastructure with straightforward instances, reserved capacity, and a path to large AI clusters.
- CoreWeave suits large-scale enterprise training on full NVIDIA HGX H100 nodes with specialist distributed infrastructure.
- Google Cloud and Amazon EC2 suit organizations keeping H100 workloads within an existing hyperscaler ecosystem, while Modal suits Python-native serverless inference and batch processing.
Production H100 applications need more than GPU capacity. They also need APIs, workers, databases, storage, CI/CD, private networking, autoscaling, secrets, logs, and governance. Northflank runs these components with H100 services and jobs on one platform, using Northflank Cloud, self-serve BYOC in your AWS, Azure, Google Cloud, Oracle, Civo, or CoreWeave account, or your existing on-premises and bare-metal Kubernetes infrastructure via BYOK.
Enterprises can retain control over networking, IAM, regional placement, data residency, and cloud billing while using SSO, RBAC, audit logs, and standardized deployment workflows. Startups and smaller engineering teams can use the same deployment model without building and operating their own Kubernetes platform.
Request NVIDIA H100 GPU capacity, get started self-serve, or book a demo for architecture, security, compliance, or migration requirements.
The accelerator may have the same name, but the hardware configuration and platform around it can produce a very different operational result.
- Workload and hardware fit: Match PCIe, SXM, or NVL configurations and single- or multi-GPU nodes to inference, batch jobs, or distributed training.
- Access and total cost: Compare on-demand, spot, reserved, and sales-assisted capacity, including CPU, memory, storage, transfer, idle time, and interruption risk.
- Operating model: Decide whether you need direct infrastructure, serverless execution, or a platform that also manages deployment, scaling, observability, databases, and security controls.
No single infrastructure or platform model is the right choice for every H100 deployment. The useful comparison is which provider fits the way your team builds and operates its workload.
The NVIDIA H100 is a Hopper architecture data-center GPU for large-model training, high-throughput inference, and high-performance computing (HPC). The right accelerator depends on model memory, precision, batch size, communication between GPUs, software support, and total job time. H200 or Blackwell GPUs may fit workloads needing more memory or newer hardware, while smaller inference services may cost less on an L4 or A100.
For specifications, purchase costs, and rental considerations, read How much does an NVIDIA H100 GPU cost?.
The table summarizes the differences that determine which platform fits a workload. Prices are as at 30 July 2026.
| Platform | Best for | Deployment | H100 GPU pricing | Pricing scope |
|---|---|---|---|---|
| Northflank | Running complete H100 applications with GPU and application infrastructure together | Managed cloud, self-serve BYOC, or existing Kubernetes via BYOK | H100 80 GB: $2.74/GPU-hour; A100, H200, B200, and more available | GPU rate includes attached CPU and RAM; billed per second |
| Runpod | Renting H100s with a choice of form factors and consumption models | Pods, Serverless, or Clusters | Runpod Secure Cloud: H100 PCIe 80 GB at $2.89/hour, H100 SXM 80 GB at $2.99/hour, and H100 NVL 94 GB at $3.19/hour. | Pod configurations include CPU and RAM; storage and deployment model affect total cost |
| Lambda | Moving from self-serve H100 nodes to reserved AI clusters | Instances, 1-Click Clusters, and Superclusters | 8× H100 SXM 80 GB: $3.99 per GPU-hour for self-serve instances. | Listed instance is an eight-GPU node; taxes may apply |
| CoreWeave | Running large, tightly coupled training workloads | Full NVIDIA HGX H100 nodes | 8× H100 node: $49.24/hour on demand; $19.51/hour spot in Europe or $19.71/hour spot in North America | Full-node pricing; calculated equivalent is about $6.16/GPU-hour on demand or $2.44–$2.46/GPU-hour spot, before other charges. |
| Google Cloud | Running H100 workloads in GKE and the Google Cloud AI stack | A3 High, A3 Mega, and A3 Edge VMs | Varies by machine type, region, and provisioning model | Storage, networking, commitments, and capacity model affect total cost |
| Amazon EC2 | Running H100 workloads in established AWS environments | One- or eight-GPU P5 instances and EC2 UltraClusters | Varies by region and purchasing model | Storage, data transfer, reservations, quotas, and adjacent services affect total cost |
| Modal | Running bursty Python inference and batch functions | Serverless functions and containers | NVIDIA H100: $3.95/hour | CPU and memory are billed separately |
These rates are not like for like. Included compute, storage, networking, interruption terms, operating-system costs, and platform services differ.
The seven platforms below address different parts of the H100 deployment problem, from complete production applications to direct GPU infrastructure and serverless functions.
Choose Northflank when your H100 workload is part of a larger production AI application, especially when you need infrastructure choice and enterprise governance without building the surrounding platform yourself.
- H100 pricing: The Northflank GPU platform lists NVIDIA H100 80 GB at $2.74/hour, including the GPU and its attached CPU and RAM. Billing is prorated to the second.
- Complete application platform: Run H100 inference, training, and jobs alongside APIs, workers, databases, storage, CI/CD, autoscaling, secrets, logs, and metrics in one workflow.
- Infrastructure choice: Use Northflank Cloud, self-serve BYOC across AWS, Google Cloud, Azure, Oracle, or CoreWeave, or eligible on-premises and bare-metal Kubernetes through BYOK.
- Enterprise governance: BYOC and BYOK retain control over networking, IAM, regions, data residency, and cloud billing. Northflank also supports SSO, RBAC, audit logs, secure runtime options, and integrations with existing registries, Vault, and DNS.
- GPU operations: The GPU documentation covers GPU-backed services and jobs, while the optimisation guide covers scaling, storage, health checks, networking, and startup commands.
Unlike providers that primarily supply GPU instances or clusters, Northflank manages the H100 workload and its surrounding application services across managed, BYOC, and BYOK infrastructure.
Choose Northflank when you need to run H100 workloads alongside the rest of your production application, or when infrastructure choice, enterprise governance, and a consistent workflow across managed cloud, BYOC, and BYOK are priorities.
The Weights engineering team used Northflank to operate nine clusters, more than 40 microservices, and over 250 concurrent GPUs without a dedicated DevOps team. This example demonstrates production GPU orchestration rather than H100-specific usage.
Request NVIDIA H100 capacity, get started (self-serve), or book a demo to discuss enterprise requirements.
Runpod suits teams that want self-serve H100 access and a choice between dedicated Pods, serverless workers, and multi-node clusters.
- H100 options and pricing: Pods list PCIe 80 GB at $2.89/hour, SXM 80 GB at $2.99/hour, and NVL 94 GB at $3.19/hour, billed per second.
- Deployment models: Pods provide dedicated instances, Serverless targets inference, and Clusters support multi-node jobs with shared storage.
- Scaling: Clusters support up to 64 GPUs with shared storage, but H100 SXM cluster pricing requires contacting Runpod’s sales team.
Choose Runpod when flexible access, form-factor choice, and Pod-level control are more important than operating the complete application stack on one platform.
Lambda suits teams that want dedicated AI infrastructure with a path from self-serve H100 instances to reserved clusters.
- Configuration and pricing: The eight-GPU H100 SXM 80 GB node is listed at $3.99 per GPU-hour, plus applicable tax.
- Cluster options: H100 clusters range from 16 to more than 2,000 GPUs. Published pricing covers terms from two weeks to one year, with longer commitments available through a sales-assisted path.
- Capacity model: Instances are self-serve and first-come, while reservations use a sales-assisted path.
Choose Lambda when dedicated AI infrastructure, full-node access, and a reservation path are the primary requirements.
CoreWeave suits organizations running large, tightly coupled training workloads on full HGX systems rather than renting a single general-purpose GPU instance.
- Configuration: The NVIDIA HGX H100 node contains eight 80 GB GPUs, 128 vCPUs, 2,048 GB RAM, and 61.44 TB local storage.
- Pricing: The node lists at $49.24/hour on demand. Spot is region-specific at $19.71/hour in North America and $19.51/hour in Europe.
- Hardware options: The catalogue also includes H200, B200, B300, and RTX PRO 6000 systems.
Choose CoreWeave when large distributed training and specialist HGX infrastructure take priority over single-GPU flexibility or a broader application platform.
Google Cloud suits organizations that want H100 capacity inside an established environment with existing data, networking, governance, GKE, and managed AI services.
- Configurations: A3 High offers one, two, four, or eight H100 80 GB GPUs. A3 Mega is a fixed eight-GPU configuration for large-scale training and serving, while A3 Edge is a fixed eight-GPU configuration designed for serving in selected regions.
- Provisioning: Smaller A3 High shapes require Spot or Flex-start; Google recommends GKE or Slurm for A3 Mega.
- Pricing: Rates vary by machine, region, and provisioning model, with quotas and capacity affecting planning.
Choose Google Cloud when native GKE, data, and AI service integration are more important than a simpler specialist-cloud workflow.
If you need to keep H100 workloads, data, networking, IAM, and cloud billing within your Google Cloud account without taking on the operational overhead of building and maintaining the surrounding application platform, Northflank BYOC for Google Cloud provides a managed layer for orchestration, deployment, scaling, CI/CD, observability, databases, and GPU workloads. This gives enterprises greater control over data residency, security, and existing Google Cloud commitments while reducing the infrastructure their platform teams must operate directly.
Amazon EC2 suits organizations that want H100 capacity inside established AWS IAM, VPC, storage, automation, governance, procurement, and managed-service workflows.
- Configurations: P5 includes one-GPU
p5.4xlargeand eight-GPUp5.48xlargeshapes using H100 80 GB GPUs. - Networking and scale: The eight-GPU shape provides EFA, GPUDirect RDMA, 3,200 Gbps networking, and 900 GB/s NVSwitch; EC2 UltraClusters extend P5 to large fleets.
- Pricing: Costs vary by region and purchase model, with storage, transfer, reservations, and quotas affecting the bill.
Choose Amazon EC2 when AWS identity, networking, purchasing, and service integration determine the architecture.
If you need to keep H100 workloads, data, networking, IAM, and cloud billing within your AWS account without taking on the operational overhead of building and maintaining the surrounding application platform, Northflank BYOC for AWS provides a managed layer for orchestration, deployment, scaling, CI/CD, observability, databases, and GPU workloads. This helps enterprises retain control over data residency, security, and existing AWS commitments while reducing the infrastructure their platform teams must operate directly.
Modal suits Python-centric teams running bursty inference, batch processing, and functions that benefit from per-second execution and scale-to-zero economics.
- Configuration: Modal supports one to eight NVIDIA H100 SXM GPUs per container.
- Pricing: Modal lists NVIDIA H100 GPU tasks at $3.95/hour. CPU and memory are billed separately.
- Execution model: Developers assign H100 GPUs to Python functions.
- Capacity behavior: Modal may automatically upgrade an H100 request to an H200. Multi-node training is currently in private beta at this time of writing.
Choose Modal for Python-based inference and batch workloads that fit a GPU-backed function model.
Match the provider to your application’s operating model rather than one hourly rate. Choose Northflank for complete production AI applications, enterprise BYOC, or BYOK; Runpod for flexible H100 access; Lambda for dedicated nodes and reservations; CoreWeave for large HGX clusters; Google Cloud or Amazon EC2 for native hyperscaler integration; and Modal for Python-based inference and batch functions. Also consider Microsoft Azure when Azure integration or confidential H100 NVL computing is a primary requirement.
For steady workloads, calculate monthly infrastructure and operational costs. For variable workloads, test availability, startup behavior, scaling, interruption risk, and the billable execution window. See How much does an NVIDIA H100 GPU cost? for a more detailed cost breakdown.
Pricing depends on the configuration and billing model. Northflank lists H100 80 GB at $2.74 per GPU-hour with its attached CPU and RAM. CoreWeave lists a full eight-GPU HGX H100 node at $49.24/hour on demand, with spot pricing of $19.71/hour in North America and $19.51/hour in Europe. These rates are not directly comparable because node size, storage, networking, interruption risk, and commitments differ. Read How much does an NVIDIA H100 GPU cost? for a fuller breakdown.
PCIe suits workloads that need single-GPU flexibility. SXM is used in tightly connected multi-GPU systems, while H100 NVL provides 94 GB of memory per GPU for memory-intensive inference. Choose based on memory, topology, software compatibility, and workload behavior.
Yes, when your workload benefits from Hopper software support, 80 GB of GPU memory, FP8, or multi-GPU infrastructure. H200 and Blackwell GPUs may suit workloads that need more memory or newer hardware, while smaller inference workloads may cost less on an L4 or A100.
Choose Northflank when inference is part of a complete production application with APIs, workers, databases, CI/CD, and observability. Runpod provides Pod and Serverless options, Google Cloud A3 Edge is designed for serving in supported regions, and Modal fits Python-based inference functions.
Lambda and Runpod fit single-node and smaller training workloads. CoreWeave fits large, tightly coupled HGX clusters. Amazon EC2 and Google Cloud fit organizations that want to keep training data, identity, networking, and managed services within an existing hyperscaler environment.
Yes, when the platform supports customer-cloud deployment. Northflank’s self-serve BYOC lets teams deploy H100 workloads in their cloud account while retaining control over infrastructure, networking, IAM, regions, data residency, and cloud billing. Read the guide to deploying GPUs in your own cloud.
- How much does an NVIDIA H100 GPU cost?: Compare H100 specifications, purchase costs, cloud pricing, and deployment options.
- Rent H100 GPU: Pricing, performance, and where to get one: Review the decisions involved in renting H100 capacity.
- Best cloud services for renting high-performance GPUs on demand: Compare a wider selection of GPU services.
- Best serverless GPU providers: Evaluate scale-to-zero GPU platforms for inference.
- Best options for BYOC in cloud computing: Compare platforms that run in your own cloud account.
- Top GPU sandboxes for AI agents: Compare governed GPU execution environments for agent workloads.

