← Back to Blog
Header image for blog post: How to run AI coding agents in GCP
Simo Aleksandrov
Published 6th October 2026

How to run AI coding agents in GCP

TL;DR: how to run AI coding agents in GCP

  • You can run AI coding agents such as Claude Code, Codex, or OpenCode in Google Cloud on a Compute Engine VM, as a Cloud Run job, on a platform you build on GKE, or in a managed environment deployed into your own GCP project. Running agents in GCP gives you control over their execution environment and lets them access private resources such as Cloud SQL and internal services in your VPC.
  • Claude Code can call Claude models through Vertex AI, with model usage billed to your GCP project. Running agents on Compute Engine yourself works for simple setups, but you take on isolation, patching, credentials, idle costs, and team access.
  • Northflank Cloud Harnesses provide the infrastructure to run coding agents in isolated cloud environments inside your own GCP project through self-serve BYOC, with support for Claude Code, Codex, Cursor, OpenCode, Pi, or bring your own agent. Harnesses also run on Northflank's managed cloud, with microVM isolation, persistent storage, configurable compute and networking, a web terminal, and SSH access.

Run AI coding agents in your own GCP project with Northflank Cloud Harnesses, or book a demo to discuss your setup.

Teams that run on Google Cloud usually want their coding agents there too. Source code already lives in repositories connected to GCP workloads, test databases run on Cloud SQL, and security teams already review access through IAM and Cloud Audit Logs. An agent running on a developer laptop sits outside all of that.

Running coding agents in your own GCP project gives you more control over their execution environment, network access, and identity permissions. The agent can work alongside your existing GCP infrastructure, authenticate with a service account, and call Claude through Vertex AI. This guide covers the ways to run AI coding agents in GCP, how to set one up on Compute Engine with Vertex AI, where that approach breaks down, and how to run coding agents in your GCP project with Northflank Cloud Harnesses.

Why run AI coding agents in GCP?

Running coding agents in your GCP project gives you more control over where they run, which credentials they have, and which resources they can access. The agent can reach Cloud SQL instances, Memorystore, and internal services over private networking in your VPC, so it can run integration tests against real dependencies without exposing them publicly. IAM service accounts define what the agent can access, and Cloud Audit Logs record the Google Cloud API calls it makes.

Compute runs on your Google Cloud bill, so it counts toward committed use discounts and appears in your existing billing reports and budgets. With Vertex AI, Claude Code can make model requests through Google Cloud, with usage billed to your project. This lets you run the coding environment and model access through your existing GCP infrastructure.

What are the ways to run coding agents in GCP?

Compute Engine VMCloud Run jobNorthflank Cloud Harnesses with BYOC on GCPSelf-built platform on GKE
Setup effortMedium for one VM, grows with each environmentMedium: container image, job configuration, networkingLow: connect your GCP project, then create HarnessesHigh: build and operate the platform yourself
IsolationTasks share the VM unless you configure separate environmentsIsolated container execution per jobIsolated microVM per HarnessDepends on the platform you build
Interactive accessSSH or IAP tunnelLimited, designed for batch runsWeb terminal and SSHWhatever you build
PersistencePersistent disk on the VMEphemeral by defaultPersistent workspace when enabledWhatever you build
Team accessOS Login or SSH keys you manageCloud console or gcloudTeammates can connect through their Northflank accountsWhatever you build
Supported agentsAny agent you installAny agent in your imageClaude Code, Codex, Cursor, OpenCode, Pi, or your own agentAny agent in your images
Ongoing maintenancePatching, tooling, cleanupImages and job definitionsNorthflank manages the Harness platform layerThe whole platform

A Compute Engine VM works for a basic setup, but every additional agent means another environment to configure and maintain. Northflank Cloud Harnesses give each coding agent its own isolated cloud environment, ready to use with the compute, storage, networking, credentials, web terminal, and SSH access it needs. You can create multiple Harnesses for parallel tasks, pause them when they're not needed, and run them in your own GCP infrastructure through BYOC. That gives you the control of your cloud account without having to build a coding agent platform yourself.

How to run a coding agent on Compute Engine

  1. Create a service account for the agent and grant it only what it needs, such as reading a secret from Secret Manager and calling Vertex AI:
gcloud iam service-accounts create coding-agent

gcloud projects add-iam-policy-binding YOUR-PROJECT-ID \
  --member="serviceAccount:coding-agent@YOUR-PROJECT-ID.iam.gserviceaccount.com" \
  --role="roles/aiplatform.user"

gcloud secrets add-iam-policy-binding github-token \
  --member="serviceAccount:coding-agent@YOUR-PROJECT-ID.iam.gserviceaccount.com" \
  --role="roles/secretmanager.secretAccessor"
  1. Create a VM that runs as that service account:
gcloud compute instances create coding-agent \
  --zone=us-central1-a \
  --machine-type=e2-standard-4 \
  --service-account=coding-agent@YOUR-PROJECT-ID.iam.gserviceaccount.com \
  --scopes=cloud-platform \
  --image-family=ubuntu-2404-lts-amd64 \
  --image-project=ubuntu-os-cloud
  1. Connect and install Git, tmux, Node.js, and the agent CLI:
gcloud compute ssh coding-agent --zone=us-central1-a

sudo apt-get update && sudo apt-get install -y git tmux
# Install Node.js using your preferred method, then:
npm install -g @anthropic-ai/claude-code
  1. Read the Git credential from Secret Manager at runtime instead of writing it to disk:
export GITHUB_TOKEN=$(gcloud secrets versions access latest --secret=github-token)
  1. Clone the repository, create a branch, and start the agent inside tmux so the terminal session remains available when you disconnect:
git clone https://github.com/your-org/your-repo.git
cd your-repo && git checkout -b task-142
tmux new -s task-142
claude
  1. Detach with Ctrl+B, then D, and reattach later with:
tmux attach -t task-142

What breaks when you run coding agents on Compute Engine yourself?

ProblemWhat happens on Compute EngineOn Northflank Cloud Harnesses in your GCP project
Isolation between tasksTasks share the VM unless you provision separate environmentsEach Harness runs in an isolated microVM
Idle costThe VM continues consuming GCP resources while it is runningPause or delete a Harness when its task is done
Patching and toolingYou patch the OS and keep agent CLIs and runtimes up to dateThe configured agent CLI comes pre-installed
CredentialsYou manage Git, model, Google Cloud, and application credentials on the VMCredentials can be configured per Harness
Team accessYou manage OS Login roles or SSH access for each developerTeammates can connect through their Northflank accounts
Scaling to more agentsMore tasks mean more VMs to create, configure, and clean upCreate additional Harnesses in the same cluster
Persistence and recoveryYou manage persistent disks and what happens when the VM restartsPersistent workspace can preserve files across restarts

Running one coding agent on Compute Engine is straightforward. The complexity comes when you need to manage multiple isolated environments, credentials, networking, persistence, and team access.

Northflank Cloud Harnesses handle that infrastructure for you, while letting the Harnesses run in your own GCP project through BYOC.

How do Northflank Cloud Harnesses run coding agents in your GCP project?

Think of a Harness as a dedicated cloud environment for your coding agent, without the infrastructure work normally required to build one.

Northflank Cloud Harnesses are cloud workspaces built for coding agents. Each Harness is configured for one agent at creation, whether Claude Code, Codex, Cursor, OpenCode, Pi, or bring your own agent, with the agent CLI pre-installed and its credentials injected as environment variables. A Harness can clone a connected Git repository and branch on start, and you work in it through the web terminal in the Northflank dashboard or over SSH from your local terminal.

image.png

With BYOC on GCP, Northflank provisions and manages a GKE cluster inside your GCP project, and your Harnesses run on it. This puts the Harness workload and its workspace inside your GCP infrastructure, while letting the agent reach Cloud SQL and internal services over private networking when your VPC routing and firewall configuration allows it. The underlying compute is billed by GCP.

Every Harness runs in an isolated microVM. When workspace persistence is enabled, files in the Harness workspace are retained across restarts. You can pause a Harness and resume it later, or delete it when the task is done. Teammates with access can connect to the same Harness through their own Northflank accounts.

For Claude Code, you can combine BYOC with Vertex AI by configuring Claude Code to use Google Cloud as its model provider. This lets Claude Code make model requests through Vertex AI while the coding environment runs in your GCP infrastructure.

Create your first Harness on Northflank, or follow How to run a coding agent in the cloud for the full Harness setup.

How to connect your GCP project to Northflank

Let your coding agent build its own environment

You don't even have to configure everything manually. With the Northflank Skill, you can give your coding agent a single command and let it create and configure the Northflank resources it needs.

  1. Connect your GCP project. In the Northflank dashboard, create a provider link to your Google Cloud project.
  2. Create a cluster. Northflank provisions a GKE cluster in your project, in the region you choose.
  3. Add node pools. Configure the node pools that will run your workloads, including machine types and zones.
  4. Deploy workloads. Your cluster is ready for projects and Harnesses.

See Integrate your Google Cloud account for the full steps and required GCP permissions. If you already run a GKE cluster, you can also import an existing cluster for Northflank to manage.

How to create a Harness in your GCP project

  1. Create a Northflank project on your GCP cluster. When you create a project, choose your GCP cluster as where its resources deploy instead of Northflank's cloud.
  2. Create a Harness in that project. Select your agent, environment size, authentication, and the repository and branch to work on. Add any other secrets the task needs, such as Vertex AI settings for Claude Code.
  3. Connect and run the agent. Open the Harness in the web terminal, or connect over SSH with the Northflank CLI, and start the agent:
northflank dev ssh --projectId your-project-id --harnessId your-harness-id
claude

The Harness runs as a microVM on your GKE cluster, inside your GCP project.

FAQ

Does my code leave my GCP project?

With Northflank BYOC on GCP, the Harness runs on a GKE cluster in your own Google Cloud project, so the repository clone and agent workspace run in your infrastructure. Northflank manages the cluster and Harness configuration. Where the agent sends model requests depends on its configuration: Claude Code with Vertex AI sends model requests through Google Cloud, while an agent using another model provider's API sends requests to that provider.

Can the agent use Claude through Vertex AI?

Yes. Claude Code supports Vertex AI as a model provider. Configure the required Google Cloud project, region, authentication, and Vertex AI permissions, then enable Vertex AI in Claude Code. See the Claude Code documentation for the current configuration.

Can the agent reach a Cloud SQL instance with a private IP?

Yes, if the network allows it. With Northflank BYOC, Harnesses run on a GKE cluster in your GCP environment, so they can reach private resources when your VPC and firewall rules permit it. Give the agent a scoped database user rather than an application or admin user.

Who pays for the compute?

With BYOC, the compute resources running in your GCP project are billed by Google Cloud, so they count toward your Google Cloud spend and commitments. Claude usage through Vertex AI is also billed to your project. See Northflank pricing for the BYOC platform pricing.

Can I use an existing GKE cluster?

Yes. Northflank can provision a new GKE cluster in your project, or you can import an existing Kubernetes cluster for Northflank to manage. See Import an existing cluster.

Share this article with your network
X