Hosted CI runners are convenient, fast to set up, and metered by the minute. They are also rented compute that does not always fit your workload. At some scale or for some specific needs, self-hosted runners — your own machines, your own network — become attractive.
The decision is more nuanced than "they are cheaper at scale." Self-hosted runners shift cost from a monthly bill to engineering time and operational risk. This post is when that shift pays off and how to do it without surprising yourself.
Why You Would Self-Host
Five real reasons:
Cost at scale. Hosted runners cost $0.005–$0.01 per Linux minute. At 100,000 minutes/month — a single mid-sized engineering team — that is $500–$1,000/month, every month. A cluster of self-hosted runners can do the same work for the cost of the underlying compute, often half that.
Network or data access. Your CI needs to reach internal services not exposed to the public internet. Database test runs, internal APIs, on-prem dependencies. Self-hosted runners on your network can reach them; hosted runners cannot.
Specific hardware. GPU access for ML workloads. ARM builds for niche targets. Large memory or disk requirements. Faster CPUs than what hosted plans offer.
Compliance. Some regulated industries forbid running build artifacts on third-party infrastructure. Self-hosted runners on your own hardware satisfy the requirement.
Faster feedback for very large workloads. A monorepo where a single CI run takes an hour on hosted runners might take 15 minutes on beefier self-hosted machines.
If none of these apply, hosted is the right call. Self-hosting adds operational burden you should avoid until you have a reason.
The Operational Cost
Self-hosting a runner is not "set up a VM and forget it." It is real, ongoing operational work:
- Runner versions drift. Each CI platform releases runner updates monthly. Out-of-date runners fail in subtle ways.
- Disk fills up. Caches, intermediate artifacts, Docker layers. Without aggressive cleanup, runners eventually fail with "no space left on device."
- Compromised runners are a risk. A bad actor with PR write access can run code on your runner. If the runner has access to internal networks or secrets, the blast radius is significant.
- Capacity management. Runners pinned at idle waste money; runners running hot create queue buildup. Auto-scaling is non-trivial.
- Updates and patches. OS patches, language runtime updates, security patches. Someone owns this.
Plan for half an engineer's time to operate a real self-hosted runner setup. That cost competes with the hosted bill you would have paid.
Ephemeral vs Persistent
The most important design decision: do runners persist between jobs, or get destroyed after each one?
Persistent runners. A long-lived VM that runs many jobs over its lifetime. Faster startup (the runner is already warm), better caching (Docker images, npm cache, Composer cache persist). Higher security risk (a malicious job can affect later jobs).
Ephemeral runners. A fresh runner for each job, destroyed after. Slower startup. No state leakage between jobs. The standard for production setups in 2026.
For untrusted code (open-source projects, forks, PRs from external contributors), ephemeral runners are the only safe choice. A persistent runner with PR write access to public repos is a credential and supply-chain disaster waiting to happen.
For trusted-only code (internal team, no external contributors), persistent runners are acceptable but most teams still prefer ephemeral for simpler reasoning.
Auto-Scaling on Kubernetes
The dominant pattern in 2026 is auto-scaling ephemeral runners on Kubernetes.
For GitHub Actions, the actions-runner-controller (ARC) project runs in Kubernetes and spawns pod-per-job runners on demand. Idle = zero pods. A queued job triggers a pod start.
# Actions Runner Controller config
apiVersion: actions.github.com/v1alpha1
kind: AutoscalingRunnerSet
metadata:
name: ci-runners
spec:
githubConfigUrl: https://github.com/yourorg
minRunners: 1
maxRunners: 50
template:
spec:
containers:
- name: runner
image: ghcr.io/actions/actions-runner:latest
GitLab Runner has a similar Kubernetes executor.
The math: you pay for the Kubernetes nodes (which can auto-scale themselves on AWS via Karpenter or Cluster Autoscaler). When CI is idle, costs are minimal. When CI is busy, capacity expands.
Sizing and Capacity
Picking how many runners to provision is the first art.
- Start with peak concurrent jobs × 1.2 as your max capacity.
- Min runners of 1–2 for instant response to the first job.
- Auto-scale aggressively up, slowly down. Burst capacity is what runners are for; idle cost is the constraint.
Watch queue time as the primary metric. If jobs are waiting more than 30 seconds for a runner during normal hours, increase max. If runners are idle 90% of the time, decrease min.
Caching Strategies
A common reason teams self-host is to control caching. Hosted runners often have generic caches; self-hosted can have local volumes that persist across jobs.
For Docker builds, a dedicated BuildKit cache backed by S3 or a local volume can cut multi-minute builds to seconds:
- uses: docker/build-push-action@v5
with:
cache-from: type=s3,region=us-east-1,bucket=our-buildkit-cache
cache-to: type=s3,region=us-east-1,bucket=our-buildkit-cache,mode=max
For dependency caches (npm, Composer, pip), mounting a persistent volume into ephemeral pods gives you persistent caching with ephemeral runners. The next pod starts fresh but reads the cached dependencies.
The caching work is often the biggest perf win from self-hosting.
Security Considerations
A self-hosted runner accepts code from your repository and executes it. The threat model has to be explicit:
- Network access. A compromised job has whatever network access the runner has. If your runner can reach the production database, a malicious PR can too. Isolate aggressively.
- Secret exposure. Secrets passed to the runner are visible to the job. Use OIDC where available to get short-lived credentials instead of long-lived secrets.
- Cross-job contamination. Persistent runners can leak state across jobs (cached secrets, malicious binaries in PATH). Ephemeral runners mostly eliminate this.
- Supply chain. The runner itself is a binary you trust. Pin versions, verify signatures, monitor for vulnerabilities.
For runners that handle PRs from external contributors, treat the runner network as untrusted. No access to internal services, no access to production secrets.
Hybrid Setups
Most mature teams end up with a hybrid:
- Hosted runners for public/open-source repos and untrusted PRs
- Self-hosted runners for internal repos, secret-bearing jobs, network-restricted tests
- Hosted macOS for occasional iOS builds (the cost of self-hosting macOS is rarely worth it)
The hybrid lets each workload run where it makes the most sense. The cost is more configuration complexity.
When Not to Self-Host
- Team size under ~10 engineers — the operational overhead outweighs the savings
- Hosted CI costs are under $500/month — not worth the work
- No engineer has the time or skills to operate the runner cluster
- Compliance and network access do not require it
For most small and medium teams, hosted CI is the right answer. Self-hosting becomes attractive in the same band where teams hit hosted-CI cost or capability limits.
A Practical Path
If you want to self-host:
- Start with one or two self-hosted runners for specific high-value workloads (the slow integration test suite, the GPU job). Keep hosted for everything else.
- Move workloads to self-hosted as they justify the cost.
- Adopt auto-scaling Kubernetes runners once you have multiple workloads sharing capacity.
- Eventually, most CI runs on self-hosted; hosted is the fallback for untrusted code.
The migration is gradual. Going all-in on self-hosted in week one is a recipe for surprised engineers and broken CI.
Looking at a CI bill that has grown faster than the team, or constrained by what hosted runners can do? We help teams design self-hosted runner setups that save money without inviting operational pain. scopeforged.com