Enterprise Platform Engineering: 90-Day Plan & Build vs Buy
The Shift from DevOps to Platform Engineering
We have spent years building and operating infrastructure for enterprises ranging from financial institutions to global retailers. Over the past twelve months, the conversations we are having with CTOs and VPs of Engineering have shifted decisively. The question is no longer "should we adopt DevOps?" but rather "how do we build an internal platform that makes our developers genuinely productive?"
This is not just industry chatter. Gartner forecasts that 80% of large software engineering organisations will have dedicated platform teams by the end of 2026. A recent survey of over 500 practitioners found that 94% view platform engineering as critical or important to their organisation's future. After deploying platform engineering initiatives across dozens of enterprise environments, we can confirm that this shift is real, urgent, and often poorly executed.
What Platform Engineering Actually Means
Platform engineering is the discipline of designing and building self-service toolchains and workflows that enable software engineering teams to deliver value without wrestling with underlying infrastructure. The output is typically an Internal Developer Platform, or IDP, that abstracts away the complexity of Kubernetes clusters, CI/CD pipelines, observability stacks, and cloud resource provisioning.
The critical distinction is intent. A DevOps team that builds shared CI/CD pipelines is doing infrastructure work. A platform engineering team treats those pipelines as a product, with defined interfaces, versioned APIs, documentation, onboarding flows, and feedback loops. When we build platforms for our clients, we apply the same product management discipline to internal tooling that you would apply to a customer-facing SaaS application.
Why Most Enterprise Platform Initiatives Stall
We see the same failure patterns repeatedly. Understanding them is the first step toward building a platform that actually delivers.
- Portal-first thinking. Too many teams start by building a shiny developer portal with Backstage or a similar tool, then struggle to connect it to meaningful backend automation. The portal becomes a dashboard that nobody uses because it does not actually reduce toil. Start with the automation layer. The interface comes last.
- No product ownership. Only 21% of organisations have a dedicated Platform Product Manager. Without someone who owns the backlog, conducts user research with internal developers, and prioritises ruthlessly, the platform becomes a dumping ground for every infrastructure wish list item.
- Mandate-driven adoption. Roughly a third of organisations try to force developers onto their platform through mandates. This works briefly and then collapses. If your platform does not create genuine pull, where developers choose to use it because it makes their lives better, you have a governance tool masquerading as a platform.
- Chronic underfunding. Nearly half of platform teams operate on less than one million dollars annually. That budget might cover a small team maintaining existing CI/CD, but it will not fund the product-grade platform that modern enterprises need. You cannot build a platform that serves hundreds of engineers on a shoestring.
The Architecture That Works
After building platforms for organisations running everything from monolithic Java applications to distributed microservices on Kubernetes, we have converged on an architecture that consistently delivers results.
Layer 1: Infrastructure Abstraction
At the foundation sits infrastructure-as-code with proper governance. Every cloud resource, from VPCs to managed databases to GPU clusters, is provisioned through Terraform or Pulumi with policy-as-code guardrails evaluated on every commit. In 2026, with AI-generated infrastructure changes accelerating the volume of modifications, GitOps enforcement is no longer a best practice. It is basic operational hygiene. If your infrastructure can be changed meaningfully outside Git, you do not have governance. You have drift.
Layer 2: Application Runtime
Kubernetes remains the standard runtime for most enterprise workloads, but the platform must abstract its complexity away from application developers. Our teams build opinionated deployment manifests, Helm charts, or custom operators that let developers specify what they need (language runtime, resource limits, scaling behaviour) without writing raw Kubernetes YAML. The developer should think in terms of their application, not in terms of pods and services.
Layer 3: Developer Workflows
This is where the platform becomes tangible to its users. Golden paths for common scenarios: spinning up a new microservice, creating a staging environment, running database migrations, triggering canary deployments. Each golden path is a paved road that encodes your organisation's best practices into an automated workflow. Developers who follow the golden path get security scanning, observability, and compliance checks for free. Those who need to deviate can, but they accept the additional responsibility.
Layer 4: Observability and Feedback
A platform without observability is flying blind. We integrate distributed tracing, structured logging, and metrics collection into the golden paths so that every application deployed through the platform is observable by default. Crucially, we also instrument the platform itself. You need to know how long it takes a developer to go from "git push" to "running in production," and you need to track that metric obsessively.
The AI Dimension
Platform engineering in 2026 has a dual mandate that did not exist two years ago. Platforms must simultaneously incorporate AI capabilities into their own tooling and provide the infrastructure for AI workloads at scale.
On the first front, we are embedding AI into platform workflows: intelligent code review, automated incident triage, natural language interfaces for infrastructure provisioning. These are not gimmicks. When a developer can type "create a staging environment for service X with a PostgreSQL database" and have the platform provision it correctly in minutes, you have eliminated hours of YAML wrangling and Slack messages to the infrastructure team.
On the second front, enterprises need their platforms to handle GPU scheduling, model serving infrastructure, training pipeline orchestration, and the unique cost dynamics of AI workloads. Traditional autoscaling based on CPU and memory does not work when your bottleneck is GPU utilisation and your cost driver is token throughput. We build platform extensions that understand these AI-specific requirements and expose them through the same self-service model that handles conventional workloads.
Measuring What Matters
The measurement crisis in platform engineering is real. Nearly 30% of organisations do not measure platform success at all. Without metrics, you cannot justify continued investment, and without investment, your platform stagnates.
We recommend tracking four categories from day one:
- Developer velocity. Time from commit to production. Time to onboard a new service. Time to provision a development environment. These measure whether the platform is actually reducing friction.
- Platform reliability. Uptime of platform services, CI/CD success rates, mean time to recover from platform incidents. Your platform is production infrastructure. Treat it accordingly.
- Adoption quality. Not just how many teams use the platform, but how they use it. Are they following golden paths? Are they working around the platform? High adoption with high workaround rates means your abstractions are wrong.
- Business impact. Deployment frequency, change failure rate, lead time for changes. These DORA metrics connect platform investment to engineering outcomes that leadership understands.
Should you build an internal developer platform or buy one?
This is the first question every enterprise engineering leader asks us, and the honest answer is that almost every platform we have delivered is a hybrid. Buying gives you a working portal and a catalogue in weeks; building gives you abstractions that match how your organisation actually ships software. The split we recommend is simple: buy the commodity layers, build the parts that encode your own rules.
- Buy or adopt open source for the portal and service catalogue (Backstage, Port, Cortex), CI runners, secret management, artifact registries and observability backends. These are solved problems and maintaining your own version is pure cost.
- Build your golden paths, environment provisioning, policy-as-code guardrails and the deployment contract between application teams and infrastructure. These carry your compliance requirements, your cloud account topology and your release process. No vendor knows them.
- Do not build a bespoke portal before you have automation worth exposing. We have inherited three Backstage installations that indexed repositories and did nothing else; each was abandoned within a year.
A useful test: if a capability would look roughly the same at any other company in your sector, buy it. If explaining it requires describing your organisation, build it.
| Platform layer | Recommended approach | Typical time to value | Why |
|---|---|---|---|
| Developer portal & service catalogue | Buy / adopt open source | 2–6 weeks | Commodity capability; maintaining a bespoke portal is pure cost |
| CI runners, artifact registry, secrets | Buy | 2–4 weeks | Solved problems with mature managed options |
| Golden paths & service templates | Build | 6–12 weeks | Encodes your release process and team topology |
| Environment provisioning | Build on managed IaC | 8–12 weeks | Depends on your cloud account topology and network model |
| Policy-as-code guardrails | Build on an open-source engine | 4–8 weeks | Compliance rules are organisation-specific |
| Observability backend | Buy | 2–4 weeks | Scale economics favour vendors; differentiation is in instrumentation |
How big should an enterprise platform team be, and what does it cost?
Across the enterprise platform programmes we have staffed, the ratio that holds is roughly one platform engineer for every twenty to thirty application engineers, with a floor of four. Below four people the team cannot run production infrastructure and build new capability at the same time; every roadmap item is eaten by on-call.
A workable initial team for an organisation of 200 engineers looks like: a Platform Product Manager, two senior infrastructure engineers, two senior software engineers who write platform services rather than YAML, and a part-time security engineer for policy-as-code review. Fully loaded, that is a seven-figure annual commitment in most Western markets — which is precisely why underfunded platform teams stall. The comparison that matters to a CFO is not the cost of the team but the cost of the friction it removes: if 200 engineers each lose four hours a week to environment provisioning, failed pipelines and deployment coordination, that is roughly twenty full-time engineers of lost capacity.
Where budget is genuinely constrained, staff seniority over headcount. A platform built by two engineers who have operated systems at scale outperforms one built by six who have not, because the abstractions are chosen better and there is less to unpick later.
What does a realistic 90-day rollout look like?
We run enterprise platform engagements on the same three-phase cadence, and it is deliberately unglamorous.
- Days 1-30 — measure and choose one path. Instrument the current state: lead time from commit to production, environment provisioning time, CI/CD failure rate, and the number of human handoffs per release. Interview fifteen to twenty engineers. Pick the single highest-friction workflow for your most common application type. Ship nothing yet.
- Days 31-60 — build the first golden path end to end. One application archetype, from repository template through pipeline, policy checks, environment provisioning and deploy, with observability wired in by default. Onboard two volunteer teams, not twenty. Fix what they complain about.
- Days 61-90 — prove the delta and fund the next phase. Re-measure the same four metrics against the baseline. Publish the numbers internally. Use them to justify the second golden path and the permanent team structure.
Teams that follow this sequence typically report lead-time reductions in the range of 40 to 70 percent on the first workflow. The number matters less than the fact that it is measured — it is the evidence that keeps the programme funded past its first year.
Common enterprise constraints we plan around
Enterprise platform work rarely fails on technology. It fails on the constraints nobody wrote down: a regulated change-approval board that requires a human sign-off before production, three cloud accounts inherited from acquisitions, an offshore delivery partner with different tooling, or a mainframe integration that cannot move. Design the platform around these from day one. Encode the change approval as a pipeline step rather than pretending it does not exist; expose the legacy integration behind a golden path so teams stop hand-rolling connections to it. A platform that ignores the organisation's real constraints is a platform that teams route around.
How do you make an enterprise platform reliable at scale?
Once dozens of teams deploy through one platform, the platform itself becomes production-critical. An outage in the pipeline or the developer portal stops every release at once. We treat it with the same discipline as a customer-facing system:
- Platform SLOs. Publish targets for pipeline availability, deploy lead time and environment provisioning time, with error budgets that decide when the platform team pauses feature work to fix reliability.
- No single control plane. Run CI runners, artifact registries and GitOps controllers redundantly across zones, and make sure a broken portal never blocks a deploy that can go straight through Git.
- Progressive delivery by default. Canary and blue-green rollouts with automatic rollback on SLO breach are built into the golden path, so reliability does not depend on each team building it themselves.
- Versioned platform changes. Templates, modules and base images are versioned and rolled out in waves. Upgrading the platform should never force every team to change on the same day.
Solving the integration challenges of enterprise platform engineering
Integration is where most enterprise platform programmes lose time. The platform has to connect to identity (Okta, Entra ID), ticketing and change management (ServiceNow, Jira), secrets stores, existing CI systems, and legacy estates that will not move to Kubernetes. Our senior engineers handle this with a few rules: wrap each integration behind one platform API instead of letting teams call it directly; adopt existing CI and identity rather than replacing them in phase one; sync change records automatically from pipeline events so audit trails come for free; and give legacy and mainframe systems their own golden path with contract tests, so they are covered by the same observability and security controls as new services.
Getting Started Without Boiling the Ocean
The biggest mistake we see is attempting to build the entire platform at once. Enterprises that succeed start small, prove value, and expand.
Begin with a single high-friction workflow, usually the path from code commit to production deployment for your most common application type. Build a golden path for that workflow. Instrument it. Measure the before and after. Use that data to justify expanding to the next workflow.
Hire or designate a Platform Product Manager before you write a single line of platform code. This person talks to developers, understands their pain points, and ensures you are building what they need rather than what the infrastructure team thinks they should want.
Finally, invest in the team. Platform engineering requires a blend of infrastructure expertise, software engineering skill, and product thinking that is genuinely rare. Senior engineers who have operated systems at scale, who understand both the developer experience and the operational reality, are the foundation of every successful platform initiative we have delivered.
Platform engineering is not a tool you install or a framework you adopt. It is an organisational capability that compounds over time. The enterprises that invest now will have a structural advantage in engineering velocity and operational resilience that their competitors will spend years trying to replicate.
Need help with your next project?
Book a free 30-minute discovery session with our senior engineers to discuss your specific challenges.