Docs

Compute Plans and Quota

Use a resource plan to choose CPU, RAM, and GPU resources for each instance of a service or job. Account billing plans determine your account limits and are separate from resource plans.

The tables below use the public Northflank API. They update from the current plan and region lists. Listed GPU support does not guarantee available capacity or account access.

Compute plans

Set billing.deploymentPlan to a deployment plan ID. For builds, set billing.buildPlan to a plan that supports build. RAM values use MiB, where 1024 MiB equals 1 GiB.

Plan IDvCPURAM (MiB)Use
nf-compute-100.1256deployment
nf-compute-200.2512deployment
nf-compute-500.51024deployment
nf-compute-100-111024deployment
nf-compute-100-212048deployment
nf-compute-100-414096deployment
nf-compute-20024096deployment
nf-compute-200-424096deployment
nf-compute-200-828192deployment
nf-compute-200-16216384deployment
nf-compute-40048192deployment
nf-compute-400-16416384deployment, build
nf-compute-800-888192deployment, build
nf-compute-800-16816384deployment, build
nf-compute-800-24824576deployment, build
nf-compute-800-32832768deployment, build
nf-compute-800-40840960deployment, build
nf-compute-1200-241224576deployment, build
nf-compute-1600-321632768deployment, build
nf-compute-2000-402040960deployment, build
nf-compute-32-13107232131072deployment
nf-compute-32-48832499712deployment

For current prices, see Northflank pricing. Query the complete public list without an API token:

curl --fail https://api.northflank.com/v1/plans

Dedicated compute, custom plans, and bare-metal

Northflank supports dedicated compute, custom plans on its managed platform (PaaS), and bare-metal servers. Bare-metal servers provide dedicated physical hardware.

Custom PaaS plans support up to 384 vCPU, 3 TB of memory, and 12 TB of NVMe storage.

If you need bare-metal or 10,000–150,000 vCPU, Book a call to discuss your requirements.

GPU plans

GPU plans bundle CPU and system RAM for the selected GPU model and count. GPU memory is separate from system RAM. Set billing.deploymentPlan to nf-gpu-<gpuType>-<gpuCount>g, replacing underscores in the plan ID with hyphens. Keep the original GPU type in deployment.gpu.configuration.gpuType.

Set the matching type and count in deployment.gpu.configuration:

{
  "billing": { "deploymentPlan": "nf-gpu-rtx-pro-6000-96-1g" },
  "deployment": {
    "gpu": {
      "enabled": true,
      "configuration": {
        "gpuType": "rtx_pro_6000-96",
        "gpuCount": 1,
        "timesliced": false
      }
    }
  }
}

The matching compute plan, GPU type, and GPU count are required. Plan errors name the matching billing.deploymentPlan. Ephemeral storage and shared memory are optional. See GPU storage for their defaults.

Use the same configuration for services and jobs. Counts apply to each instance. Timeslicing, which shares a GPU between workloads, is available in your own cloud.

Plan IDGPU typeGPUs per instanceMemory per GPU (GiB)Regions
nf-gpu-a100-40-1ga100-40140asia-northeast, asia-southeast, europe-west-netherlands, us-central
nf-gpu-a100-40-2ga100-40240asia-northeast, asia-southeast, europe-west-netherlands, us-central
nf-gpu-a100-40-4ga100-40440asia-northeast, asia-southeast, europe-west-netherlands, us-central
nf-gpu-a100-40-8ga100-40840asia-northeast, asia-southeast, europe-west-netherlands, us-central
nf-gpu-a100-80-1ga100-80180asia-southeast, europe-west-netherlands, us-central, us-east1
nf-gpu-a100-80-2ga100-80280asia-southeast, europe-west-netherlands, us-central, us-east1
nf-gpu-a100-80-4ga100-80480asia-southeast, europe-west-netherlands, us-central, us-east1
nf-gpu-a100-80-8ga100-80880asia-southeast, europe-west-netherlands, us-central, us-east1
nf-gpu-b200-180-8gb200-1808180asia-northeast, asia-southeast, europe-west-netherlands, us-central, us-east1
nf-gpu-h100-80-1gh100-80180asia-southeast, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west
nf-gpu-h100-80-2gh100-80280asia-southeast, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west
nf-gpu-h100-80-4gh100-80480asia-southeast, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west
nf-gpu-h100-80-8gh100-80880asia-northeast, asia-southeast, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west
nf-gpu-h200-141-8gh200-1418141europe-west-netherlands, us-central, us-east1, us-west
nf-gpu-l4-24-1gl4-24124asia-northeast, asia-southeast, europe-west, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west
nf-gpu-l4-24-2gl4-24224asia-northeast, asia-southeast, europe-west, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west
nf-gpu-l4-24-4gl4-24424asia-northeast, asia-southeast, europe-west, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west
nf-gpu-l4-24-8gl4-24824asia-northeast, asia-southeast, europe-west, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west
nf-gpu-rtx-pro-6000-96-1grtx_pro_6000-96196asia-south-delhi, us-central, us-east-ohio, us-west
nf-gpu-rtx-pro-6000-96-2grtx_pro_6000-96296asia-south-delhi, us-central, us-east-ohio, us-west
nf-gpu-rtx-pro-6000-96-4grtx_pro_6000-96496asia-south-delhi, us-central, us-east-ohio, us-west
nf-gpu-rtx-pro-6000-96-8grtx_pro_6000-96896asia-south-delhi, europe-west, us-central, us-east-ohio, us-west

The table derives GPU plan IDs from each region's supported types and counts. The /v1/plans endpoint lists compute plans. Use /v1/regions to discover GPU types and counts.

GPU regions

Create your project in a region that supports the GPU type and count you need. The project region cannot change after creation. Counts can differ between regions for the same GPU type.

RegionAPI referenceGlobal regionGPU types and counts per instance
Asia Northeastasia-northeastAsia Pacific
b200-180: 8
h100-80: 8
a100-40: 1, 2, 4, 8
l4-24: 1, 2, 4, 8
Asia South Delhiasia-south-delhiAsia Pacific
rtx_pro_6000-96: 1, 2, 4, 8
Asia Southeastasia-southeastAsia Pacific
b200-180: 8
h100-80: 1, 2, 4, 8
a100-80: 1, 2, 4, 8
a100-40: 1, 2, 4, 8
l4-24: 1, 2, 4, 8
Europe Westeurope-westEMEA
l4-24: 1, 2, 4, 8
rtx_pro_6000-96: 8
Europe West Frankfurteurope-west-frankfurtEMEA
l4-24: 1, 2, 4, 8
h100-80: 1, 2, 4, 8
Europe West Netherlandseurope-west-netherlandsEMEA
h200-141: 8
h100-80: 1, 2, 4, 8
a100-80: 1, 2, 4, 8
a100-40: 1, 2, 4, 8
b200-180: 8
l4-24: 1, 2, 4, 8
US Centralus-centralAmericas
h200-141: 8
h100-80: 1, 2, 4, 8
a100-80: 1, 2, 4, 8
a100-40: 1, 2, 4, 8
l4-24: 1, 2, 4, 8
rtx_pro_6000-96: 1, 2, 4, 8
b200-180: 8
US East Ohious-east-ohioAmericas
rtx_pro_6000-96: 1, 2, 4, 8
US East1us-east1Americas
b200-180: 8
h200-141: 8
h100-80: 1, 2, 4, 8
a100-80: 1, 2, 4, 8
l4-24: 1, 2, 4, 8
US Westus-westAmericas
h200-141: 8
h100-80: 1, 2, 4, 8
l4-24: 1, 2, 4, 8
rtx_pro_6000-96: 1, 2, 4, 8

Query the region list without an API token:

curl --fail https://api.northflank.com/v1/regions

See GPU sandbox examples for one or eight H100 or RTX PRO 6000 GPUs.

Quotas and GPU access

Your account and project quotas limit resource counts and resource sizes. A plan listed above does not grant access beyond those limits. Free projects do not support GPU workloads.

Default quotas

These are selected default quotas for managed cloud workloads. Account and project limits can differ. Free projects have separate limits.

QuotaDefaultScope
CPU8 vCPUPer instance
Memory16 GiBPer instance
Ephemeral storage2 GiBPer instance
Shared memory (SHM)64 MiBPer instance
Projects100Per account or team
Services500Across projects in an account or team
Jobs100Across projects in an account or team
Service instances10Per service
Addon replicas3Per addon
Concurrent job runs10Per job

Concurrent job runs are separate executions of the same job. Free jobs allow one active run at a time.

The resource sizes above apply to standard CPU workloads. Managed GPU workloads use the selected GPU plan's CPU, memory, and storage allocations. For a quota increase, contact support@northflank.com.

For GPU access requirements, see GPU credits and access. Contact support@northflank.com if your account cannot use a supported GPU configuration.

GPU storage

Ephemeral storage is temporary disk space that does not survive container replacement. Managed GPU workloads use a fixed allocation for each GPU type and count. This allocation differs from the CPU ephemeral storage quota.

deployment.storage.ephemeralStorage.storageSize is optional for managed GPU workloads. The API sets it to the allocation per GPU multiplied by deployment.gpu.configuration.gpuCount, even if you supply another value. Changes to the GPU type or count recalculate this allocation. Storage values use MiB.

GPU typeEphemeral storage per GPU (MiB)
h100-80512000
rtx_pro_6000-961228800
a100-40, a100-80, l4-24256000
h200-141, b200-1801280000

For example, H100 with gpuCount: 8 receives storageSize: 4096000. RTX PRO 6000 with gpuCount: 8 receives storageSize: 9830400. These allocations come from the GPU type, not GPU memory or shared memory.

The GPU examples omit ephemeral storage and let the API select the allocation. CPU workloads and workloads in your own cloud use their existing storage defaults and limits.

Shared memory (deployment.storage.shmSize) uses system RAM and is separate from ephemeral storage and GPU memory. This field is optional. For managed GPU workloads, the API defaults it to system RAM per GPU multiplied by gpuCount. Smaller explicit values are preserved. Values above that limit are capped. Changes to GPU type or count reapply this rule to the existing value. CPU workloads and workloads in your own cloud keep the 64 MiB default. The GPU examples explicitly request 32 GiB. A persistent volume has its own size and quota.

Troubleshoot API requests

If the API reports a GPU plan mismatch or says the plan is only available for specific GPU types, set billing.deploymentPlan to nf-gpu-<gpuType>-<gpuCount>g, replacing underscores in the plan ID with hyphens. Include deployment.gpu.enabled: true and the matching type and count under deployment.gpu.configuration. Use the exact plan named in the error. Make sure that the project region supports that type and count.

Older API versions can require explicit GPU storage. If they report exceeds maximum ephemeral storage capacity or does not meet the minimum ephemeral storage, supply the exact allocation from the table. Both smaller and larger values fail on those versions. Older errors label these numbers as MB; the API field uses MiB.

For an older API version, the storage fragment for H100 with eight GPUs is:

{
  "deployment": {
    "storage": {
      "ephemeralStorage": { "storageSize": 4096000 }
    }
  }
}

If the API reports a runtime ephemeral storage allowance error for a GPU request, first check GPU enablement, configuration, and the matching plan. Do not reduce GPU storage to the standard CPU quota. For a CPU request, reduce storage to your account or project allowance, or request a quota increase.

If a correctly configured GPU request still fails, send support the error, project region, GPU type and count, plan ID, and storage size. Omit your API token.

Resource count quotas and per-instance resource limits are separate. Increasing a jobs count quota does not increase CPU, RAM, or storage limits.

© 2026 Northflank Ltd. All rights reserved.

northflank.com / Terms / Privacy / feedback@northflank.com