# Compute Plans and Quota

Use a resource plan to choose CPU, RAM, and GPU resources for each instance of a service or job. Account billing plans determine your account limits and are separate from resource plans.

The tables below use the public Northflank API. They update from the current plan and region lists. Listed GPU support does not guarantee available capacity or account access.

## Compute plans

Set `billing.deploymentPlan` to a deployment plan ID. For builds, set `billing.buildPlan` to a plan that supports `build`. RAM values use MiB, where 1024 MiB equals 1 GiB.

| Plan ID | vCPU | RAM (MiB) | Use |
| --- | --- | --- | --- |
| `nf-compute-10` | 0.1 | 256 | deployment |
| `nf-compute-20` | 0.2 | 512 | deployment |
| `nf-compute-50` | 0.5 | 1024 | deployment |
| `nf-compute-100-1` | 1 | 1024 | deployment |
| `nf-compute-100-2` | 1 | 2048 | deployment |
| `nf-compute-100-4` | 1 | 4096 | deployment |
| `nf-compute-200` | 2 | 4096 | deployment |
| `nf-compute-200-4` | 2 | 4096 | deployment |
| `nf-compute-200-8` | 2 | 8192 | deployment |
| `nf-compute-200-16` | 2 | 16384 | deployment |
| `nf-compute-400` | 4 | 8192 | deployment |
| `nf-compute-400-16` | 4 | 16384 | deployment, build |
| `nf-compute-800-8` | 8 | 8192 | deployment, build |
| `nf-compute-800-16` | 8 | 16384 | deployment, build |
| `nf-compute-800-24` | 8 | 24576 | deployment, build |
| `nf-compute-800-32` | 8 | 32768 | deployment, build |
| `nf-compute-800-40` | 8 | 40960 | deployment, build |
| `nf-compute-1200-24` | 12 | 24576 | deployment, build |
| `nf-compute-1600-32` | 16 | 32768 | deployment, build |
| `nf-compute-2000-40` | 20 | 40960 | deployment, build |
| `nf-compute-32-131072` | 32 | 131072 | deployment |
| `nf-compute-32-488` | 32 | 499712 | deployment |

For current prices, see [Northflank pricing](https://northflank.com/pricing). Query the complete public list without an API token:

```shell
curl --fail https://api.northflank.com/v1/plans
```

### Dedicated compute, custom plans, and bare-metal

Northflank supports dedicated compute, custom plans on its managed platform (PaaS), and bare-metal servers. Bare-metal servers provide dedicated physical hardware.

Custom PaaS plans support up to 384 vCPU, 3 TB of memory, and 12 TB of NVMe storage.

If you need bare-metal or 10,000–150,000 vCPU, [Book a call](https://cal.com/team/northflank/northflank-intro) to discuss your requirements.

## GPU plans

GPU plans bundle CPU and system RAM for the selected GPU model and count. GPU memory is separate from system RAM. Set `billing.deploymentPlan` to `nf-gpu-<gpuType>-<gpuCount>g`, replacing underscores in the plan ID with hyphens. Keep the original GPU type in `deployment.gpu.configuration.gpuType`.

Set the matching type and count in `deployment.gpu.configuration`:

```json
{
  "billing": { "deploymentPlan": "nf-gpu-rtx-pro-6000-96-1g" },
  "deployment": {
    "gpu": {
      "enabled": true,
      "configuration": {
        "gpuType": "rtx_pro_6000-96",
        "gpuCount": 1,
        "timesliced": false
      }
    }
  }
}
```

The matching compute plan, GPU type, and GPU count are required. Plan errors name the matching `billing.deploymentPlan`. Ephemeral storage and shared memory are optional. See [GPU storage](#gpu-storage) for their defaults.

Use the same configuration for services and jobs. Counts apply to each instance. Timeslicing, which shares a GPU between workloads, is available in your own cloud.

| Plan ID | GPU type | GPUs per instance | Memory per GPU (GiB) | Regions |
| --- | --- | --- | --- | --- |
| `nf-gpu-a100-40-1g` | `a100-40` | 1 | 40 | asia-northeast, asia-southeast, europe-west-netherlands, us-central |
| `nf-gpu-a100-40-2g` | `a100-40` | 2 | 40 | asia-northeast, asia-southeast, europe-west-netherlands, us-central |
| `nf-gpu-a100-40-4g` | `a100-40` | 4 | 40 | asia-northeast, asia-southeast, europe-west-netherlands, us-central |
| `nf-gpu-a100-40-8g` | `a100-40` | 8 | 40 | asia-northeast, asia-southeast, europe-west-netherlands, us-central |
| `nf-gpu-a100-80-1g` | `a100-80` | 1 | 80 | asia-southeast, europe-west-netherlands, us-central, us-east1 |
| `nf-gpu-a100-80-2g` | `a100-80` | 2 | 80 | asia-southeast, europe-west-netherlands, us-central, us-east1 |
| `nf-gpu-a100-80-4g` | `a100-80` | 4 | 80 | asia-southeast, europe-west-netherlands, us-central, us-east1 |
| `nf-gpu-a100-80-8g` | `a100-80` | 8 | 80 | asia-southeast, europe-west-netherlands, us-central, us-east1 |
| `nf-gpu-b200-180-8g` | `b200-180` | 8 | 180 | asia-northeast, asia-southeast, europe-west-netherlands, us-central, us-east1 |
| `nf-gpu-h100-80-1g` | `h100-80` | 1 | 80 | asia-southeast, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west |
| `nf-gpu-h100-80-2g` | `h100-80` | 2 | 80 | asia-southeast, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west |
| `nf-gpu-h100-80-4g` | `h100-80` | 4 | 80 | asia-southeast, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west |
| `nf-gpu-h100-80-8g` | `h100-80` | 8 | 80 | asia-northeast, asia-southeast, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west |
| `nf-gpu-h200-141-8g` | `h200-141` | 8 | 141 | europe-west-netherlands, us-central, us-east1, us-west |
| `nf-gpu-l4-24-1g` | `l4-24` | 1 | 24 | asia-northeast, asia-southeast, europe-west, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west |
| `nf-gpu-l4-24-2g` | `l4-24` | 2 | 24 | asia-northeast, asia-southeast, europe-west, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west |
| `nf-gpu-l4-24-4g` | `l4-24` | 4 | 24 | asia-northeast, asia-southeast, europe-west, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west |
| `nf-gpu-l4-24-8g` | `l4-24` | 8 | 24 | asia-northeast, asia-southeast, europe-west, europe-west-frankfurt, europe-west-netherlands, us-central, us-east1, us-west |
| `nf-gpu-rtx-pro-6000-96-1g` | `rtx_pro_6000-96` | 1 | 96 | asia-south-delhi, us-central, us-east-ohio, us-west |
| `nf-gpu-rtx-pro-6000-96-2g` | `rtx_pro_6000-96` | 2 | 96 | asia-south-delhi, us-central, us-east-ohio, us-west |
| `nf-gpu-rtx-pro-6000-96-4g` | `rtx_pro_6000-96` | 4 | 96 | asia-south-delhi, us-central, us-east-ohio, us-west |
| `nf-gpu-rtx-pro-6000-96-8g` | `rtx_pro_6000-96` | 8 | 96 | asia-south-delhi, europe-west, us-central, us-east-ohio, us-west |

The table derives GPU plan IDs from each region's supported types and counts. The `/v1/plans` endpoint lists compute plans. Use `/v1/regions` to discover GPU types and counts.

## GPU regions

Create your project in a region that supports the GPU type and count you need. The project region cannot change after creation. Counts can differ between regions for the same GPU type.

| Region | API reference | Global region | GPU types and counts per instance |
| --- | --- | --- | --- |
| Asia Northeast | `asia-northeast` | Asia Pacific | `b200-180`: 8`h100-80`: 8`a100-40`: 1, 2, 4, 8`l4-24`: 1, 2, 4, 8 |
| Asia South Delhi | `asia-south-delhi` | Asia Pacific | `rtx_pro_6000-96`: 1, 2, 4, 8 |
| Asia Southeast | `asia-southeast` | Asia Pacific | `b200-180`: 8`h100-80`: 1, 2, 4, 8`a100-80`: 1, 2, 4, 8`a100-40`: 1, 2, 4, 8`l4-24`: 1, 2, 4, 8 |
| Europe West | `europe-west` | EMEA | `l4-24`: 1, 2, 4, 8`rtx_pro_6000-96`: 8 |
| Europe West Frankfurt | `europe-west-frankfurt` | EMEA | `l4-24`: 1, 2, 4, 8`h100-80`: 1, 2, 4, 8 |
| Europe West Netherlands | `europe-west-netherlands` | EMEA | `h200-141`: 8`h100-80`: 1, 2, 4, 8`a100-80`: 1, 2, 4, 8`a100-40`: 1, 2, 4, 8`b200-180`: 8`l4-24`: 1, 2, 4, 8 |
| US Central | `us-central` | Americas | `h200-141`: 8`h100-80`: 1, 2, 4, 8`a100-80`: 1, 2, 4, 8`a100-40`: 1, 2, 4, 8`l4-24`: 1, 2, 4, 8`rtx_pro_6000-96`: 1, 2, 4, 8`b200-180`: 8 |
| US East Ohio | `us-east-ohio` | Americas | `rtx_pro_6000-96`: 1, 2, 4, 8 |
| US East1 | `us-east1` | Americas | `b200-180`: 8`h200-141`: 8`h100-80`: 1, 2, 4, 8`a100-80`: 1, 2, 4, 8`l4-24`: 1, 2, 4, 8 |
| US West | `us-west` | Americas | `h200-141`: 8`h100-80`: 1, 2, 4, 8`l4-24`: 1, 2, 4, 8`rtx_pro_6000-96`: 1, 2, 4, 8 |

Query the region list without an API token:

```shell
curl --fail https://api.northflank.com/v1/regions
```

See [GPU sandbox examples](/docs/v1/application/sandboxes/examples#create-a-gpu-sandbox) for one or eight H100 or RTX PRO 6000 GPUs.

## Quotas and GPU access

Your account and project quotas limit resource counts and resource sizes. A plan listed above does not grant access beyond those limits. Free projects do not support GPU workloads.

### Default quotas

These are selected default quotas for managed cloud workloads. Account and project limits can differ. Free projects have separate limits.

| Quota | Default | Scope |
| --- | --- | --- |
| CPU | 8 vCPU | Per instance |
| Memory | 16 GiB | Per instance |
| Ephemeral storage | 2 GiB | Per instance |
| Shared memory (SHM) | 64 MiB | Per instance |
| Projects | 100 | Per account or team |
| Services | 500 | Across projects in an account or team |
| Jobs | 100 | Across projects in an account or team |
| Service instances | 10 | Per service |
| Addon replicas | 3 | Per addon |
| Concurrent job runs | 10 | Per job |

Concurrent job runs are separate executions of the same job. Free jobs allow one active run at a time.

The resource sizes above apply to standard CPU workloads. Managed GPU workloads use the selected GPU plan's CPU, memory, and storage allocations. For a quota increase, contact [support@northflank.com](mailto:support@northflank.com).

For GPU access requirements, see [GPU credits and access](/docs/v1/application/billing/credits#gpu-access). Contact [support@northflank.com](mailto:support@northflank.com) if your account cannot use a supported GPU configuration.

### GPU storage

Ephemeral storage is temporary disk space that does not survive container replacement. Managed GPU workloads use a fixed allocation for each GPU type and count. This allocation differs from the CPU ephemeral storage quota.

`deployment.storage.ephemeralStorage.storageSize` is optional for managed GPU workloads. The API sets it to the allocation per GPU multiplied by `deployment.gpu.configuration.gpuCount`, even if you supply another value. Changes to the GPU type or count recalculate this allocation. Storage values use MiB.

| GPU type | Ephemeral storage per GPU (MiB) |
| --- | --- |
| `h100-80` | 512000 |
| `rtx_pro_6000-96` | 1228800 |
| `a100-40`, `a100-80`, `l4-24` | 256000 |
| `h200-141`, `b200-180` | 1280000 |

For example, H100 with `gpuCount: 8` receives `storageSize: 4096000`. RTX PRO 6000 with `gpuCount: 8` receives `storageSize: 9830400`. These allocations come from the GPU type, not GPU memory or shared memory.

The [GPU examples](/docs/v1/application/sandboxes/examples#create-a-gpu-sandbox) omit ephemeral storage and let the API select the allocation. CPU workloads and workloads in your own cloud use their existing storage defaults and limits.

Shared memory (`deployment.storage.shmSize`) uses system RAM and is separate from ephemeral storage and GPU memory. This field is optional. For managed GPU workloads, the API defaults it to system RAM per GPU multiplied by `gpuCount`. Smaller explicit values are preserved. Values above that limit are capped. Changes to GPU type or count reapply this rule to the existing value. CPU workloads and workloads in your own cloud keep the 64 MiB default. The GPU examples explicitly request 32 GiB. A persistent volume has its own size and quota.

### Troubleshoot API requests

If the API reports a GPU plan mismatch or says the plan is only available for specific GPU types, set `billing.deploymentPlan` to `nf-gpu-<gpuType>-<gpuCount>g`, replacing underscores in the plan ID with hyphens. Include `deployment.gpu.enabled: true` and the matching type and count under `deployment.gpu.configuration`. Use the exact plan named in the error. Make sure that the project region supports that type and count.

Older API versions can require explicit GPU storage. If they report `exceeds maximum ephemeral storage capacity` or `does not meet the minimum ephemeral storage`, supply the exact allocation from the table. Both smaller and larger values fail on those versions. Older errors label these numbers as `MB`; the API field uses MiB.

For an older API version, the storage fragment for H100 with eight GPUs is:

```json
{
  "deployment": {
    "storage": {
      "ephemeralStorage": { "storageSize": 4096000 }
    }
  }
}
```

If the API reports a runtime ephemeral storage allowance error for a GPU request, first check GPU enablement, configuration, and the matching plan. Do not reduce GPU storage to the standard CPU quota. For a CPU request, reduce storage to your account or project allowance, or request a quota increase.

If a correctly configured GPU request still fails, send support the error, project region, GPU type and count, plan ID, and storage size. Omit your API token.

Resource count quotas and per-instance resource limits are separate. Increasing a jobs count quota does not increase CPU, RAM, or storage limits.
