When every team selects its own models, agents, tools and integrations, adoption can outpace the organisation’s ability to secure and support it. Credentials proliferate, the same capabilities are rebuilt repeatedly and nobody has a complete inventory of which agents can act on which systems.
Platform engineering provides a route from experimentation to controlled scale. A successful platform does not force every use case into one framework. It offers reusable foundations, safe execution environments and paved roads for common patterns while preserving room for justified variation.
This course guides participants through the architecture and operating model of an internal agentic development platform. It combines developer experience with identity, policy, evaluation, observability and cost governance.
Learning Outcomes
Upon completion of this course, participants will be able to:
- Define the responsibilities and boundaries of an agentic development platform
- Design shared access to models, tools, data and execution environments
- Create registries and lifecycle controls for agents and MCP servers
- Provide secure sandboxes and isolated development workspaces
- Integrate evaluation, observability and policy into delivery pipelines
- Design golden paths without creating rigid framework lock-in
- Establish ownership, support and cost-allocation models
- Produce a phased platform roadmap for enterprise adoption
Course Outline
From Individual Tools to a Shared Platform
- Adoption stages from personal assistants to enterprise agent portfolios
- Risks of duplicated infrastructure and unmanaged agent sprawl
- Internal developer platforms versus agentic development platforms
- Product thinking for platform capabilities
- Defining platform customers, use cases and service boundaries
- Choosing what to centralise and what to federate
Reference Architecture
- Developer interfaces, APIs, gateways and runtime services
- Model access, tool access and knowledge services
- Identity, policy, secrets and network controls
- Evaluation, observability and cost telemetry
- Registries for agents, tools and reusable components
- Multi-account and multi-environment isolation patterns
Model and Inference Gateway
- Standard interfaces across proprietary and open models
- Approved-model catalogues and use restrictions
- Routing by task, risk, latency and budget
- Rate limits, quotas and workload prioritisation
- Logging metadata without retaining unnecessary content
- Resilience to provider, region and model-version changes
Agent and Tool Registries
- Registering owners, purpose, versions and dependencies
- Discovering approved agents and MCP capabilities
- Capability descriptions, risk ratings and service levels
- Preventing duplicate or abandoned integrations
- Vulnerability, deprecation and retirement workflows
- Supporting A2A discovery within trusted boundaries
Secure Runtimes and Workspaces
- Local, remote and ephemeral agent execution
- Sandboxed shell, browser and code environments
- Network egress and filesystem policy
- Workload identity and short-lived credentials
- Tenant, team and data-domain isolation
- Capacity, concurrency and workload scheduling
Golden Paths and Developer Experience
- Templates for common agent and coding workflows
- Approved libraries, tool contracts and instrumentation
- Self-service environments with policy built in
- Repository instructions and context services
- Documentation, examples and support channels
- Escape hatches and the process for justified exceptions
Evaluation and Delivery Pipelines
- Versioning prompts, tools, policies and datasets
- Standard offline and pre-production evaluation stages
- Security, permission and end-to-end workflow tests
- Promotion criteria and human approval
- Canary, shadow, rollback and emergency suspension
- Evidence packs generated as part of delivery
Observability and Operations
- Distributed traces across models, agents and tools
- Quality, reliability, latency and cost dashboards
- Platform service objectives and team-level indicators
- Incident response across shared and product-owned components
- Detecting unused, unhealthy or overly privileged agents
- Feeding production outcomes back into platform standards
Governance and Operating Model
- Responsibilities of platform, product, security and risk teams
- Onboarding and classification of agent use cases
- Policy-as-code and central guardrails
- Funding, showback and chargeback options
- Community contribution and reusable capability ownership
- Measuring adoption through outcomes rather than tool counts
Practical Capstone
- Map requirements for a multi-team agent platform
- Design the reference architecture and trust boundaries
- Define a golden path for one representative use case
- Add registry, policy, evaluation and observability controls
- Create a phased rollout and operating model
- Present trade-offs to technical and governance stakeholders
Intended audience
This course is intended for platform engineers, cloud architects, AI infrastructure teams, developer-experience leaders, security architects and senior engineering managers responsible for organisation-wide AI capabilities. It assumes participants are planning beyond isolated proof-of-concept deployments.
Prerequisites
Those attending this course should meet the following:
- Strong understanding of cloud and platform architecture
- Familiarity with CI/CD, containers, identity and infrastructure as code
- Basic understanding of coding agents, MCP and production AI agents
- Experience supporting multiple engineering or data teams
