

How to run AI coding agents in GCP
- You can run AI coding agents such as Claude Code, Codex, or OpenCode in Google Cloud on a Compute Engine VM, as a Cloud Run job, on a platform you build on GKE, or in a managed environment deployed into your own GCP project. Running agents in GCP gives you control over their execution environment and lets them access private resources such as Cloud SQL and internal services in your VPC.
- Claude Code can call Claude models through Vertex AI, with model usage billed to your GCP project. Running agents on Compute Engine yourself works for simple setups, but you take on isolation, patching, credentials, idle costs, and team access.
- Northflank Cloud Harnesses provide the infrastructure to run coding agents in isolated cloud environments inside your own GCP project through self-serve BYOC, with support for Claude Code, Codex, Cursor, OpenCode, Pi, or bring your own agent. Harnesses also run on Northflank's managed cloud, with microVM isolation, persistent storage, configurable compute and networking, a web terminal, and SSH access.
Run AI coding agents in your own GCP project with Northflank Cloud Harnesses, or book a demo to discuss your setup.
Teams that run on Google Cloud usually want their coding agents there too. Source code already lives in repositories connected to GCP workloads, test databases run on Cloud SQL, and security teams already review access through IAM and Cloud Audit Logs. An agent running on a developer laptop sits outside all of that.
Running coding agents in your own GCP project gives you more control over their execution environment, network access, and identity permissions. The agent can work alongside your existing GCP infrastructure, authenticate with a service account, and call Claude through Vertex AI. This guide covers the ways to run AI coding agents in GCP, how to set one up on Compute Engine with Vertex AI, where that approach breaks down, and how to run coding agents in your GCP project with Northflank Cloud Harnesses.
Running coding agents in your GCP project gives you more control over where they run, which credentials they have, and which resources they can access. The agent can reach Cloud SQL instances, Memorystore, and internal services over private networking in your VPC, so it can run integration tests against real dependencies without exposing them publicly. IAM service accounts define what the agent can access, and Cloud Audit Logs record the Google Cloud API calls it makes.
Compute runs on your Google Cloud bill, so it counts toward committed use discounts and appears in your existing billing reports and budgets. With Vertex AI, Claude Code can make model requests through Google Cloud, with usage billed to your project. This lets you run the coding environment and model access through your existing GCP infrastructure.
| Compute Engine VM | Cloud Run job | Northflank Cloud Harnesses with BYOC on GCP | Self-built platform on GKE | |
|---|---|---|---|---|
| Setup effort | Medium for one VM, grows with each environment | Medium: container image, job configuration, networking | Low: connect your GCP project, then create Harnesses | High: build and operate the platform yourself |
| Isolation | Tasks share the VM unless you configure separate environments | Isolated container execution per job | Isolated microVM per Harness | Depends on the platform you build |
| Interactive access | SSH or IAP tunnel | Limited, designed for batch runs | Web terminal and SSH | Whatever you build |
| Persistence | Persistent disk on the VM | Ephemeral by default | Persistent workspace when enabled | Whatever you build |
| Team access | OS Login or SSH keys you manage | Cloud console or gcloud | Teammates can connect through their Northflank accounts | Whatever you build |
| Supported agents | Any agent you install | Any agent in your image | Claude Code, Codex, Cursor, OpenCode, Pi, or your own agent | Any agent in your images |
| Ongoing maintenance | Patching, tooling, cleanup | Images and job definitions | Northflank manages the Harness platform layer | The whole platform |
A Compute Engine VM works for a basic setup, but every additional agent means another environment to configure and maintain. Northflank Cloud Harnesses give each coding agent its own isolated cloud environment, ready to use with the compute, storage, networking, credentials, web terminal, and SSH access it needs. You can create multiple Harnesses for parallel tasks, pause them when they're not needed, and run them in your own GCP infrastructure through BYOC. That gives you the control of your cloud account without having to build a coding agent platform yourself.
- Create a service account for the agent and grant it only what it needs, such as reading a secret from Secret Manager and calling Vertex AI:
gcloud iam service-accounts create coding-agent
gcloud projects add-iam-policy-binding YOUR-PROJECT-ID \
--member="serviceAccount:coding-agent@YOUR-PROJECT-ID.iam.gserviceaccount.com" \
--role="roles/aiplatform.user"
gcloud secrets add-iam-policy-binding github-token \
--member="serviceAccount:coding-agent@YOUR-PROJECT-ID.iam.gserviceaccount.com" \
--role="roles/secretmanager.secretAccessor"
- Create a VM that runs as that service account:
gcloud compute instances create coding-agent \
--zone=us-central1-a \
--machine-type=e2-standard-4 \
--service-account=coding-agent@YOUR-PROJECT-ID.iam.gserviceaccount.com \
--scopes=cloud-platform \
--image-family=ubuntu-2404-lts-amd64 \
--image-project=ubuntu-os-cloud
- Connect and install Git, tmux, Node.js, and the agent CLI:
gcloud compute ssh coding-agent --zone=us-central1-a
sudo apt-get update && sudo apt-get install -y git tmux
# Install Node.js using your preferred method, then:
npm install -g @anthropic-ai/claude-code
- Read the Git credential from Secret Manager at runtime instead of writing it to disk:
export GITHUB_TOKEN=$(gcloud secrets versions access latest --secret=github-token)
- Clone the repository, create a branch, and start the agent inside tmux so the terminal session remains available when you disconnect:
git clone https://github.com/your-org/your-repo.git
cd your-repo && git checkout -b task-142
tmux new -s task-142
claude
- Detach with
Ctrl+B, thenD, and reattach later with:
tmux attach -t task-142
| Problem | What happens on Compute Engine | On Northflank Cloud Harnesses in your GCP project |
|---|---|---|
| Isolation between tasks | Tasks share the VM unless you provision separate environments | Each Harness runs in an isolated microVM |
| Idle cost | The VM continues consuming GCP resources while it is running | Pause or delete a Harness when its task is done |
| Patching and tooling | You patch the OS and keep agent CLIs and runtimes up to date | The configured agent CLI comes pre-installed |
| Credentials | You manage Git, model, Google Cloud, and application credentials on the VM | Credentials can be configured per Harness |
| Team access | You manage OS Login roles or SSH access for each developer | Teammates can connect through their Northflank accounts |
| Scaling to more agents | More tasks mean more VMs to create, configure, and clean up | Create additional Harnesses in the same cluster |
| Persistence and recovery | You manage persistent disks and what happens when the VM restarts | Persistent workspace can preserve files across restarts |
Running one coding agent on Compute Engine is straightforward. The complexity comes when you need to manage multiple isolated environments, credentials, networking, persistence, and team access.
Northflank Cloud Harnesses handle that infrastructure for you, while letting the Harnesses run in your own GCP project through BYOC.
Think of a Harness as a dedicated cloud environment for your coding agent, without the infrastructure work normally required to build one.
Northflank Cloud Harnesses are cloud workspaces built for coding agents. Each Harness is configured for one agent at creation, whether Claude Code, Codex, Cursor, OpenCode, Pi, or bring your own agent, with the agent CLI pre-installed and its credentials injected as environment variables. A Harness can clone a connected Git repository and branch on start, and you work in it through the web terminal in the Northflank dashboard or over SSH from your local terminal.

With BYOC on GCP, Northflank provisions and manages a GKE cluster inside your GCP project, and your Harnesses run on it. This puts the Harness workload and its workspace inside your GCP infrastructure, while letting the agent reach Cloud SQL and internal services over private networking when your VPC routing and firewall configuration allows it. The underlying compute is billed by GCP.
Every Harness runs in an isolated microVM. When workspace persistence is enabled, files in the Harness workspace are retained across restarts. You can pause a Harness and resume it later, or delete it when the task is done. Teammates with access can connect to the same Harness through their own Northflank accounts.
For Claude Code, you can combine BYOC with Vertex AI by configuring Claude Code to use Google Cloud as its model provider. This lets Claude Code make model requests through Vertex AI while the coding environment runs in your GCP infrastructure.
Create your first Harness on Northflank, or follow How to run a coding agent in the cloud for the full Harness setup.
Let your coding agent build its own environment
You don't even have to configure everything manually. With the Northflank Skill, you can give your coding agent a single command and let it create and configure the Northflank resources it needs.
- Connect your GCP project. In the Northflank dashboard, create a provider link to your Google Cloud project.
- Create a cluster. Northflank provisions a GKE cluster in your project, in the region you choose.
- Add node pools. Configure the node pools that will run your workloads, including machine types and zones.
- Deploy workloads. Your cluster is ready for projects and Harnesses.
See Integrate your Google Cloud account for the full steps and required GCP permissions. If you already run a GKE cluster, you can also import an existing cluster for Northflank to manage.
- Create a Northflank project on your GCP cluster. When you create a project, choose your GCP cluster as where its resources deploy instead of Northflank's cloud.
- Create a Harness in that project. Select your agent, environment size, authentication, and the repository and branch to work on. Add any other secrets the task needs, such as Vertex AI settings for Claude Code.
- Connect and run the agent. Open the Harness in the web terminal, or connect over SSH with the Northflank CLI, and start the agent:
northflank dev ssh --projectId your-project-id --harnessId your-harness-id
claude
The Harness runs as a microVM on your GKE cluster, inside your GCP project.
With Northflank BYOC on GCP, the Harness runs on a GKE cluster in your own Google Cloud project, so the repository clone and agent workspace run in your infrastructure. Northflank manages the cluster and Harness configuration. Where the agent sends model requests depends on its configuration: Claude Code with Vertex AI sends model requests through Google Cloud, while an agent using another model provider's API sends requests to that provider.
Yes. Claude Code supports Vertex AI as a model provider. Configure the required Google Cloud project, region, authentication, and Vertex AI permissions, then enable Vertex AI in Claude Code. See the Claude Code documentation for the current configuration.
Yes, if the network allows it. With Northflank BYOC, Harnesses run on a GKE cluster in your GCP environment, so they can reach private resources when your VPC and firewall rules permit it. Give the agent a scoped database user rather than an application or admin user.
With BYOC, the compute resources running in your GCP project are billed by Google Cloud, so they count toward your Google Cloud spend and commitments. Claude usage through Vertex AI is also billed to your project. See Northflank pricing for the BYOC platform pricing.
Yes. Northflank can provision a new GKE cluster in your project, or you can import an existing Kubernetes cluster for Northflank to manage. See Import an existing cluster.



