← Back to Blog
Header image for blog post: Best NVIDIA H100 cloud providers for AI inference in 2026
Deborah Emeni
Published 31st July 2026

Best NVIDIA H100 cloud providers for AI inference in 2026

Renting NVIDIA H100 capacity is only one part of running production AI training and inference. This article compares application platforms, specialist GPU infrastructure, hyperscalers, and serverless execution across pricing, topology, infrastructure control, scaling, and production workload fit.

TL;DR: Best NVIDIA H100 GPU cloud providers

  • Northflank suits engineers, startups, and enterprises deploying H100 workloads alongside APIs, workers, jobs, databases, CI/CD, autoscaling, and observability on managed infrastructure or through self-serve BYOC.
  • Runpod suits teams that want flexible access to H100 PCIe, SXM, and NVL GPUs through Pods, Serverless, and Clusters.
  • Lambda suits teams that want dedicated H100 infrastructure with straightforward instances, reserved capacity, and a path to large AI clusters.
  • CoreWeave suits large-scale enterprise training on full NVIDIA HGX H100 nodes with specialist distributed infrastructure.
  • Google Cloud and Amazon EC2 suit organizations keeping H100 workloads within an existing hyperscaler ecosystem, while Modal suits Python-native serverless inference and batch processing.

Production H100 applications need more than GPU capacity. They also need APIs, workers, databases, storage, CI/CD, private networking, autoscaling, secrets, logs, and governance. Northflank runs these components with H100 services and jobs on one platform, using Northflank Cloud, self-serve BYOC in your AWS, Azure, Google Cloud, Oracle, Civo, or CoreWeave account, or your existing on-premises and bare-metal Kubernetes infrastructure via BYOK.

Enterprises can retain control over networking, IAM, regional placement, data residency, and cloud billing while using SSO, RBAC, audit logs, and standardized deployment workflows. Startups and smaller engineering teams can use the same deployment model without building and operating their own Kubernetes platform.

Request NVIDIA H100 GPU capacity, get started self-serve, or book a demo for architecture, security, compliance, or migration requirements.

What to look for in an NVIDIA H100 cloud provider

The accelerator may have the same name, but the hardware configuration and platform around it can produce a very different operational result.

  • Workload and hardware fit: Match PCIe, SXM, or NVL configurations and single- or multi-GPU nodes to inference, batch jobs, or distributed training.
  • Access and total cost: Compare on-demand, spot, reserved, and sales-assisted capacity, including CPU, memory, storage, transfer, idle time, and interruption risk.
  • Operating model: Decide whether you need direct infrastructure, serverless execution, or a platform that also manages deployment, scaling, observability, databases, and security controls.

No single infrastructure or platform model is the right choice for every H100 deployment. The useful comparison is which provider fits the way your team builds and operates its workload.

Is the NVIDIA H100 the right GPU for your workload?

The NVIDIA H100 is a Hopper architecture data-center GPU for large-model training, high-throughput inference, and high-performance computing (HPC). The right accelerator depends on model memory, precision, batch size, communication between GPUs, software support, and total job time. H200 or Blackwell GPUs may fit workloads needing more memory or newer hardware, while smaller inference services may cost less on an L4 or A100.

For specifications, purchase costs, and rental considerations, read How much does an NVIDIA H100 GPU cost?.

NVIDIA H100 cloud providers compared

The table summarizes the differences that determine which platform fits a workload. Prices are as at 30 July 2026.

PlatformBest forDeploymentH100 GPU pricingPricing scope
NorthflankRunning complete H100 applications with GPU and application infrastructure togetherManaged cloud, self-serve BYOC, or existing Kubernetes via BYOKH100 80 GB: $2.74/GPU-hour; A100, H200, B200, and more availableGPU rate includes attached CPU and RAM; billed per second
RunpodRenting H100s with a choice of form factors and consumption modelsPods, Serverless, or ClustersRunpod Secure Cloud: H100 PCIe 80 GB at $2.89/hour, H100 SXM 80 GB at $2.99/hour, and H100 NVL 94 GB at $3.19/hour.Pod configurations include CPU and RAM; storage and deployment model affect total cost
LambdaMoving from self-serve H100 nodes to reserved AI clustersInstances, 1-Click Clusters, and Superclusters8× H100 SXM 80 GB: $3.99 per GPU-hour for self-serve instances.Listed instance is an eight-GPU node; taxes may apply
CoreWeaveRunning large, tightly coupled training workloadsFull NVIDIA HGX H100 nodes8× H100 node: $49.24/hour on demand; $19.51/hour spot in Europe or $19.71/hour spot in North AmericaFull-node pricing; calculated equivalent is about $6.16/GPU-hour on demand or $2.44–$2.46/GPU-hour spot, before other charges.
Google CloudRunning H100 workloads in GKE and the Google Cloud AI stackA3 High, A3 Mega, and A3 Edge VMsVaries by machine type, region, and provisioning modelStorage, networking, commitments, and capacity model affect total cost
Amazon EC2Running H100 workloads in established AWS environmentsOne- or eight-GPU P5 instances and EC2 UltraClustersVaries by region and purchasing modelStorage, data transfer, reservations, quotas, and adjacent services affect total cost
ModalRunning bursty Python inference and batch functionsServerless functions and containersNVIDIA H100: $3.95/hourCPU and memory are billed separately

These rates are not like for like. Included compute, storage, networking, interruption terms, operating-system costs, and platform services differ.

What are the best cloud providers for NVIDIA H100 GPUs?

The seven platforms below address different parts of the H100 deployment problem, from complete production applications to direct GPU infrastructure and serverless functions.

1. Northflank

Choose Northflank when your H100 workload is part of a larger production AI application, especially when you need infrastructure choice and enterprise governance without building the surrounding platform yourself.

  • H100 pricing: The Northflank GPU platform lists NVIDIA H100 80 GB at $2.74/hour, including the GPU and its attached CPU and RAM. Billing is prorated to the second.
  • Complete application platform: Run H100 inference, training, and jobs alongside APIs, workers, databases, storage, CI/CD, autoscaling, secrets, logs, and metrics in one workflow.
  • Infrastructure choice: Use Northflank Cloud, self-serve BYOC across AWS, Google Cloud, Azure, Oracle, or CoreWeave, or eligible on-premises and bare-metal Kubernetes through BYOK.
  • Enterprise governance: BYOC and BYOK retain control over networking, IAM, regions, data residency, and cloud billing. Northflank also supports SSO, RBAC, audit logs, secure runtime options, and integrations with existing registries, Vault, and DNS.
  • GPU operations: The GPU documentation covers GPU-backed services and jobs, while the optimisation guide covers scaling, storage, health checks, networking, and startup commands.

Unlike providers that primarily supply GPU instances or clusters, Northflank manages the H100 workload and its surrounding application services across managed, BYOC, and BYOK infrastructure.

Choose Northflank when you need to run H100 workloads alongside the rest of your production application, or when infrastructure choice, enterprise governance, and a consistent workflow across managed cloud, BYOC, and BYOK are priorities.

The Weights engineering team used Northflank to operate nine clusters, more than 40 microservices, and over 250 concurrent GPUs without a dedicated DevOps team. This example demonstrates production GPU orchestration rather than H100-specific usage.

Request NVIDIA H100 capacity, get started (self-serve), or book a demo to discuss enterprise requirements.

2. Runpod

Runpod suits teams that want self-serve H100 access and a choice between dedicated Pods, serverless workers, and multi-node clusters.

  • H100 options and pricing: Pods list PCIe 80 GB at $2.89/hour, SXM 80 GB at $2.99/hour, and NVL 94 GB at $3.19/hour, billed per second.
  • Deployment models: Pods provide dedicated instances, Serverless targets inference, and Clusters support multi-node jobs with shared storage.
  • Scaling: Clusters support up to 64 GPUs with shared storage, but H100 SXM cluster pricing requires contacting Runpod’s sales team.

Choose Runpod when flexible access, form-factor choice, and Pod-level control are more important than operating the complete application stack on one platform.

3. Lambda

Lambda suits teams that want dedicated AI infrastructure with a path from self-serve H100 instances to reserved clusters.

  • Configuration and pricing: The eight-GPU H100 SXM 80 GB node is listed at $3.99 per GPU-hour, plus applicable tax.
  • Cluster options: H100 clusters range from 16 to more than 2,000 GPUs. Published pricing covers terms from two weeks to one year, with longer commitments available through a sales-assisted path.
  • Capacity model: Instances are self-serve and first-come, while reservations use a sales-assisted path.

Choose Lambda when dedicated AI infrastructure, full-node access, and a reservation path are the primary requirements.

4. CoreWeave

CoreWeave suits organizations running large, tightly coupled training workloads on full HGX systems rather than renting a single general-purpose GPU instance.

  • Configuration: The NVIDIA HGX H100 node contains eight 80 GB GPUs, 128 vCPUs, 2,048 GB RAM, and 61.44 TB local storage.
  • Pricing: The node lists at $49.24/hour on demand. Spot is region-specific at $19.71/hour in North America and $19.51/hour in Europe.
  • Hardware options: The catalogue also includes H200, B200, B300, and RTX PRO 6000 systems.

Choose CoreWeave when large distributed training and specialist HGX infrastructure take priority over single-GPU flexibility or a broader application platform.

5. Google Cloud

Google Cloud suits organizations that want H100 capacity inside an established environment with existing data, networking, governance, GKE, and managed AI services.

  • Configurations: A3 High offers one, two, four, or eight H100 80 GB GPUs. A3 Mega is a fixed eight-GPU configuration for large-scale training and serving, while A3 Edge is a fixed eight-GPU configuration designed for serving in selected regions.
  • Provisioning: Smaller A3 High shapes require Spot or Flex-start; Google recommends GKE or Slurm for A3 Mega.
  • Pricing: Rates vary by machine, region, and provisioning model, with quotas and capacity affecting planning.

Choose Google Cloud when native GKE, data, and AI service integration are more important than a simpler specialist-cloud workflow.

If you need to keep H100 workloads, data, networking, IAM, and cloud billing within your Google Cloud account without taking on the operational overhead of building and maintaining the surrounding application platform, Northflank BYOC for Google Cloud provides a managed layer for orchestration, deployment, scaling, CI/CD, observability, databases, and GPU workloads. This gives enterprises greater control over data residency, security, and existing Google Cloud commitments while reducing the infrastructure their platform teams must operate directly.

6. Amazon EC2

Amazon EC2 suits organizations that want H100 capacity inside established AWS IAM, VPC, storage, automation, governance, procurement, and managed-service workflows.

  • Configurations: P5 includes one-GPU p5.4xlarge and eight-GPU p5.48xlarge shapes using H100 80 GB GPUs.
  • Networking and scale: The eight-GPU shape provides EFA, GPUDirect RDMA, 3,200 Gbps networking, and 900 GB/s NVSwitch; EC2 UltraClusters extend P5 to large fleets.
  • Pricing: Costs vary by region and purchase model, with storage, transfer, reservations, and quotas affecting the bill.

Choose Amazon EC2 when AWS identity, networking, purchasing, and service integration determine the architecture.

If you need to keep H100 workloads, data, networking, IAM, and cloud billing within your AWS account without taking on the operational overhead of building and maintaining the surrounding application platform, Northflank BYOC for AWS provides a managed layer for orchestration, deployment, scaling, CI/CD, observability, databases, and GPU workloads. This helps enterprises retain control over data residency, security, and existing AWS commitments while reducing the infrastructure their platform teams must operate directly.

7. Modal

Modal suits Python-centric teams running bursty inference, batch processing, and functions that benefit from per-second execution and scale-to-zero economics.

  • Configuration: Modal supports one to eight NVIDIA H100 SXM GPUs per container.
  • Pricing: Modal lists NVIDIA H100 GPU tasks at $3.95/hour. CPU and memory are billed separately.
  • Execution model: Developers assign H100 GPUs to Python functions.
  • Capacity behavior: Modal may automatically upgrade an H100 request to an H200. Multi-node training is currently in private beta at this time of writing.

Choose Modal for Python-based inference and batch workloads that fit a GPU-backed function model.

How should you choose an NVIDIA H100 cloud provider?

Match the provider to your application’s operating model rather than one hourly rate. Choose Northflank for complete production AI applications, enterprise BYOC, or BYOK; Runpod for flexible H100 access; Lambda for dedicated nodes and reservations; CoreWeave for large HGX clusters; Google Cloud or Amazon EC2 for native hyperscaler integration; and Modal for Python-based inference and batch functions. Also consider Microsoft Azure when Azure integration or confidential H100 NVL computing is a primary requirement.

For steady workloads, calculate monthly infrastructure and operational costs. For variable workloads, test availability, startup behavior, scaling, interruption risk, and the billable execution window. See How much does an NVIDIA H100 GPU cost? for a more detailed cost breakdown.

Frequently asked questions about NVIDIA H100 cloud providers

How much does it cost to rent an NVIDIA H100 in the cloud?

Pricing depends on the configuration and billing model. Northflank lists H100 80 GB at $2.74 per GPU-hour with its attached CPU and RAM. CoreWeave lists a full eight-GPU HGX H100 node at $49.24/hour on demand, with spot pricing of $19.71/hour in North America and $19.51/hour in Europe. These rates are not directly comparable because node size, storage, networking, interruption risk, and commitments differ. Read How much does an NVIDIA H100 GPU cost? for a fuller breakdown.

Should you choose H100 PCIe, SXM, or NVL?

PCIe suits workloads that need single-GPU flexibility. SXM is used in tightly connected multi-GPU systems, while H100 NVL provides 94 GB of memory per GPU for memory-intensive inference. Choose based on memory, topology, software compatibility, and workload behavior.

Is an H100 still worth renting in 2026?

Yes, when your workload benefits from Hopper software support, 80 GB of GPU memory, FP8, or multi-GPU infrastructure. H200 and Blackwell GPUs may suit workloads that need more memory or newer hardware, while smaller inference workloads may cost less on an L4 or A100.

Which H100 cloud provider is best for inference?

Choose Northflank when inference is part of a complete production application with APIs, workers, databases, CI/CD, and observability. Runpod provides Pod and Serverless options, Google Cloud A3 Edge is designed for serving in supported regions, and Modal fits Python-based inference functions.

Which H100 cloud provider is best for model training?

Lambda and Runpod fit single-node and smaller training workloads. CoreWeave fits large, tightly coupled HGX clusters. Amazon EC2 and Google Cloud fit organizations that want to keep training data, identity, networking, and managed services within an existing hyperscaler environment.

Can enterprises run H100 workloads in their own cloud account?

Yes, when the platform supports customer-cloud deployment. Northflank’s self-serve BYOC lets teams deploy H100 workloads in their cloud account while retaining control over infrastructure, networking, IAM, regions, data residency, and cloud billing. Read the guide to deploying GPUs in your own cloud.

Share this article with your network
X