Azure Red Hat OpenShift (ARO) Guide: Architecture, Use Cases, and Best Practices
Kubernetes has become the de facto standard for container orchestration. It provides powerful capabilities for scheduling, scaling, and managing containerized workloads. But running Kubernetes in production—especially at enterprise scale—is operationally complex. Managing the control plane, configuring networking, securing the cluster, handling upgrades, and integrating with existing identity and governance systems requires specialized expertise that many organizations struggle to maintain.
This is where enterprise Kubernetes platforms like OpenShift add value. OpenShift builds on Kubernetes with additional capabilities: integrated CI/CD pipelines, built-in monitoring and logging, a developer-friendly web console, security context constraints, operators, and a curated ecosystem of tools and services. It provides a more complete platform experience out of the box.
Azure Red Hat OpenShift (ARO) brings this enterprise platform to Azure as a fully managed service. ARO is a fully managed, enterprise-grade Kubernetes platform jointly developed, operated, and supported by Red Hat and Microsoft. It allows you to deploy fully managed OpenShift clusters without the complexity of building, maintaining, or securing the underlying infrastructure.
The joint engineering and support model is a key differentiator: a dedicated Site Reliability Engineering (SRE) team from Red Hat and Microsoft works together to ensure high availability and resilience of your clusters. This means you get a seamless experience with integrated support from both vendors.
ARO is different from:
- Self-managed OpenShift on Azure VMs – You manage the entire OpenShift lifecycle yourself
- Azure Kubernetes Service (AKS) – Vanilla Kubernetes without the OpenShift platform layer
- Azure Container Apps – A higher-level abstraction for containerized applications without Kubernetes control
- Azure Red Hat Enterprise Linux – A Linux distribution, not a container platform
ARO provides an enterprise Kubernetes platform that combines Red Hat's OpenShift capabilities with Azure's infrastructure and services, while significantly reducing the operational burden of managing the underlying platform.
What Is Azure Red Hat OpenShift?
Azure Red Hat OpenShift (ARO) is a fully managed OpenShift service running on Azure that provides an enterprise Kubernetes platform while reducing the operational burden of managing the underlying OpenShift control plane and infrastructure.
The conceptual stack
Application Layer
↓
OpenShift Platform Layer (Developer tools, Operators, CI/CD, Monitoring)
↓
Kubernetes Layer (Container orchestration)
↓
Azure Infrastructure Layer (VMs, networking, storage)
↓
Azure Physical Infrastructure
What OpenShift adds to Kubernetes
OpenShift is not merely "Kubernetes with a UI." It adds several enterprise-grade capabilities on top of Kubernetes:
- OpenShift Web Console – A comprehensive web-based interface for both administrators and developers
- OpenShift CLI (
oc) – An enhanced CLI that extendskubectlwith OpenShift-specific commands - Projects – OpenShift's abstraction over Kubernetes namespaces with additional security and resource controls
- Routes – OpenShift's native ingress solution for exposing services externally
- Security Context Constraints (SCCs) – Fine-grained control over pod security permissions
- Operators – A framework for automating the lifecycle of complex applications
- OpenShift GitOps and Pipelines – Built-in CI/CD capabilities based on ArgoCD and Tekton
- Integrated container registry – A built-in registry for storing container images
- Integrated monitoring – Prometheus-based monitoring stack with Thanos for long-term storage
ARO vs. Kubernetes
While Kubernetes is considered CaaS (Container as a Service), OpenShift and ARO fall under the category PaaS (Platform as a Service). Unlike basic Kubernetes, OpenShift includes pre-integrated components: container management, automation, networking, CI/CD, monitoring, registry, and authentication—all tested together.
ARO vs. self-managed OpenShift
OpenShift offers two primary deployment models:
| Aspect | Self-Managed OpenShift | Azure Red Hat OpenShift (ARO) |
|---|---|---|
| Management | Organization manages installation, updates, and management | Fully managed by Red Hat and Microsoft |
| Infrastructure | Organization manages underlying infrastructure | Azure infrastructure managed by Microsoft |
| Support | Red Hat support | Joint Microsoft/Red Hat support |
| Billing | Direct Red Hat subscription + infrastructure | Billed through Azure subscription |
| Deployment | On any supported infrastructure | Exclusive to Azure |
Azure Red Hat OpenShift at a Glance
| Capability | Purpose |
|---|---|
| Managed OpenShift | Run enterprise OpenShift without managing the entire platform yourself |
| Kubernetes | Container orchestration foundation |
| OpenShift Console | Web-based administration and developer experience |
| Projects | Namespace-oriented application isolation with enhanced security |
| Routes | Expose applications externally with built-in TLS |
| Operators | Manage platform and application components |
| Azure networking | Integrate clusters with Azure VNets and subnets |
| Microsoft Entra ID | Identity integration for authentication |
| Azure Monitor | Monitoring and observability integration via Cluster Logging Forwarder and Cluster Observability Operator |
| Azure Container Registry | Private container image storage and management |
| Azure Storage | Persistent application data via PVs and PVCs |
| Azure Key Vault | Secrets and security integration |
| Autoscaling | Scale workloads with Horizontal Pod Autoscaler |
| Availability | Support resilient production cluster architectures with availability zones |
| OpenShift APIs and CLI | Automation and developer workflows |
| GPU Support | Support for NVIDIA H100 and H200 GPU-based Azure VM SKUs for AI/ML workloads |
ARO Architecture
Major layers
Azure Subscription
↓
Resource Group
↓
Virtual Network (with subnets)
↓
ARO Cluster
├── Control Plane (master nodes, managed by Azure/Red Hat)
├── Infrastructure Nodes (optional, customer-configurable)
└── Worker Nodes (customer-managed)
↓
OpenShift Workloads
├── Projects / Namespaces
├── Deployments
├── Pods
├── Services
├── Routes
└── Operators
Core components
ARO is built on Azure infrastructure services, including virtual machines, network security groups, and storage accounts, all deployed directly into your Azure subscription.
Operating System: ARO runs on Red Hat Enterprise Linux CoreOS (RHCOS), providing a secure, immutable OS optimized for running containers. RHCOS is designed specifically for container workloads with automated updates and minimal attack surface.
Control plane: The OpenShift control plane includes the API server, etcd (the cluster's key-value store), scheduler, and controller manager. In ARO, the control plane is fully managed by the joint Microsoft/Red Hat SRE team. You don't have direct access to control plane nodes or etcd.
Infrastructure nodes: ARO supports adding infrastructure nodes to host platform workloads like Ingress controllers, the registry, and cluster monitoring. This helps with larger clusters that have resource contention between user workloads and infrastructure workloads such as Prometheus.
Worker nodes: Worker nodes run your application workloads. You choose the VM size, number of nodes, and scaling behavior.
Responsibility model
| Component | Managed By |
|---|---|
| OpenShift control plane | Microsoft + Red Hat |
| Worker nodes (VMs) | Customer (configuration, scaling) |
| Node OS (RHCOS) | Automated updates by platform |
| Cluster networking | Customer (VNet configuration) |
| OpenShift API and console | Microsoft + Red Hat |
| Application workloads | Customer |
| Storage configuration | Customer |
| Identity integration | Customer |
| Monitoring configuration | Customer |
ARO Cluster Design
Key design decisions
Region – Choose Azure regions that support ARO. Cluster creation requires at least 44 vCPUs to create and run an OpenShift cluster.
Availability zones – Deploy ARO clusters across availability zones where supported to increase resilience.
Virtual network and subnets – ARO requires a virtual network with two empty subnets: one for master (control plane) nodes and one for worker nodes. You can create a new VNet or use an existing one. Each ARO cluster should use separate or dedicated subnets to avoid potential conflicts.
Public vs. private clusters:
| Aspect | Public Cluster | Private Cluster |
|---|---|---|
| API server visibility | Public endpoint | Private endpoint only |
| Ingress visibility | Public | Private |
| Egress | Default: LoadBalancer with internet egress | Restricted; requires VNet-integrated paths |
| Best for | Sandbox, development | Production, regulated workloads |
Opt for a public cluster only in situations like a "sandbox cluster" or where establishing a private method for console and API access is not feasible or desired.
Private cluster considerations: In a private ARO cluster, the control plane nodes don't have public internet access. The Microsoft.ContainerRegistry service endpoint on the master subnet is required so that private control-plane nodes can reach Azure Container Registry over the virtual network without using public IP connectivity. For private clusters, configuring the service endpoint on the master subnet is a prerequisite so the control plane can reliably access ACR without relying on public internet egress.
Multi-cluster enterprise strategy
For enterprise environments, consider:
- Development clusters – Smaller, public (or limited private access), lower cost
- Test clusters – Medium size, representative of production
- Production clusters – Private, multi-zone, high availability
- Multi-cluster management – Use OpenShift GitOps or fleet management for consistency
When to use separate clusters vs. projects
| Factor | Separate Clusters | Separate Projects |
|---|---|---|
| Isolation | Strongest (network, security, failure) | Namespace-level isolation |
| Cost | Higher (multiple clusters) | Lower (shared cluster) |
| Management | More complex | Simpler |
| Compliance | Different compliance requirements | Same compliance boundary |
| Team autonomy | Full cluster control | Project-level control |
Networking
Virtual network requirements
ARO requires a virtual network with two empty subnets:
- Master subnet – For control plane nodes (minimum /23 recommended)
- Worker subnet – For worker nodes (minimum /23 recommended)
The subnets must be empty when creating the cluster; they cannot contain existing resources.
Egress configuration
By default, public and private clusters have --outbound-type defined to LoadBalancer, meaning all clusters have open egress to the internet through the public load balancer.
To change the default behavior and restrict internet egress, set --outbound-type to UserDefinedRouting during cluster creation and set up traffic to run through a firewall solution, such as Azure Firewall or Azure NAT Gateway.
Egress options:
- NAT Gateway – Replaces routes to go through Azure NAT Gateway for egress instead of the LoadBalancer
- Azure Firewall – Routes egress traffic through Azure Firewall with granular rule control
Private clusters and ACR access
For private clusters, the Microsoft.ContainerRegistry service endpoint on the master and worker subnets provides a direct, VNet-integrated path to ACR. This is required because private control-plane and worker nodes don't have public internet access and need to pull container images for platform components and workloads.
Service endpoints and VNet encryption
ARO version 4.18+ supports installing clusters with virtual network encryption. In this version, the dependency on service endpoints has been removed, and new clusters won't create service endpoints on the VNet.
DNS and custom domains
By default, ARO uses self-signed certificates for all routes created on *.apps.<random>.<location>.aroapp.io. Many organizations want to use their own custom domains for applications. ARO supports custom domain configuration for routes.
Network security
- Use Network Security Groups (NSGs) to control traffic at the subnet level
- Use Azure Private Link to access ARO cluster API endpoints and Kubernetes LoadBalancer-type services
- Configure private endpoints for supported Azure resources to establish private access points
- ARO service supports deployment to customer virtual networks
Identity and Access Management
Identity architecture
ARO supports multiple identity providers, with Microsoft Entra ID (formerly Azure Active Directory) being the primary recommended approach.
Microsoft Entra ID
↓
OpenShift Authentication (OAuth)
↓
OpenShift RBAC
↓
Project
↓
Application
Microsoft Entra ID integration
ARO can be configured to use Microsoft Entra ID as an OpenID Connect (OIDC) identity provider. The integration involves:
- Registering an application in Azure AD for authentication
- Configuring the application registration to include optional claims (email, preferred_username) and group claims in tokens
- Configuring the ARO cluster to use Azure AD as the identity provider
- Granting permissions to individual users or groups
Group claims: OpenShift 4.10+ supports OpenID Connect group claim functionality, allowing an identity provider to provide a user's group membership for use within OpenShift. This enables mapping Azure AD security groups to OpenShift roles.
Azure RBAC vs. OpenShift RBAC vs. Kubernetes RBAC
| Aspect | Azure RBAC | OpenShift RBAC | Kubernetes RBAC |
|---|---|---|---|
| Scope | Azure resource management | OpenShift cluster resources | Kubernetes API resources |
| Purpose | Control who can manage ARO resources | Control access to OpenShift resources | Control access to Kubernetes resources |
| Users | Azure AD users and service principals | OpenShift users and service accounts | Service accounts |
| Integration | Azure-native | Entra ID + OpenShift | OpenShift-native |
Azure RBAC alone does not replace OpenShift RBAC. You need both: Azure RBAC controls who can create and manage ARO clusters, while OpenShift RBAC controls who can deploy and manage applications within the cluster.
Managed Identities and Workload Identity
ARO supports deploying clusters using Azure Managed Identities instead of service principals, enabling Workload Identity for platform operators.
Benefits of Managed Identity with Workload Identity:
- Enhanced Security: No service principal secrets to manage; no service principal secrets stored in the cluster
- Streamlined Operations: Automatic credential rotation via Azure Managed Identity; no manual secret rotation required
- Workload Identity: Platform operators use federated credentials with Azure
- Compliance: Meets security requirements for secret-free authentication; aligns with Azure security best practices
Microsoft Entra Workload ID is available in OpenShift clusters configured to use short-term credentials, starting with ARO version 4.16 and later when originally created with managed identities.
Least privilege
Apply least-privilege principles:
- Use separate service accounts for different workloads
- Grant only the permissions required for each application
- Regularly review and audit permissions
- Use OpenShift RBAC to limit access to projects and resources
Security Architecture
Security by default
ARO enforces security best practices by default, with automated updates, integrated monitoring, and compliance controls, making it suitable for running sensitive or regulated workloads. ARO provides stricter security defaults, including Security Context Constraints (SCCs).
Security Context Constraints (SCCs)
SCCs are OpenShift's mechanism for controlling pod-level security permissions. They govern:
- Whether pods can run as root
- Which capabilities are available
- Access to host resources
- Filesystem permissions
Unlike vanilla Kubernetes, OpenShift applies restrictive SCCs by default, requiring explicit opt-in for privileged operations.
Network policies
Use OpenShift network policies to control traffic between pods and services. Network policies provide:
- Pod-to-pod communication control
- Namespace-level isolation
- Service segmentation
Container image security
- Scan images for vulnerabilities before deployment
- Use trusted registries (ACR recommended)
- Use image signing where applicable
- Avoid
latesttags in production
Secrets management
- Use OpenShift Secrets for sensitive configuration
- Integrate with Azure Key Vault for enterprise secret management
- Never store secrets in container images or source code
- Rotate secrets regularly
Audit logging
- Enable audit logging for both Azure and OpenShift
- Monitor for suspicious activity
- Use Azure Monitor for centralized log collection
OpenShift Projects and Namespaces
Projects as OpenShift namespaces
OpenShift Projects are an abstraction over Kubernetes namespaces with additional features:
- Resource quotas – Limit CPU, memory, and storage usage
- Limit ranges – Set default resource requests and limits
- RBAC – Project-scoped roles and bindings
- Network policies – Per-project network isolation
Organizing with projects
Projects can be used to organize:
- Teams – Each team gets its own project(s)
- Applications – Each application has a dedicated project
- Environments – Development, test, and production projects
- Tenants – Multi-tenant SaaS applications
Resource quotas
Resource quotas prevent any single project from consuming all cluster resources:
- CPU and memory limits
- Persistent volume claim limits
- Object count limits (pods, services, etc.)
When projects are not enough
Projects provide logical isolation but not physical isolation. For stronger isolation:
- Use separate node pools with taints and tolerations
- Use separate clusters for different compliance requirements
- Use network policies for traffic segmentation
Applications and Workloads
Workload types
ARO supports all standard Kubernetes workload types:
- Deployments – Stateless applications with rolling updates
- StatefulSets – Stateful applications with persistent storage
- Jobs – One-off batch processing
- CronJobs – Scheduled batch processing
- DaemonSets – Node-level agents and daemons
Application lifecycle
Container Image (from ACR or other registry)
↓
Deployment (defines desired state)
↓
ReplicaSet (maintains replica count)
↓
Pod (running container instance)
↓
Service (stable network endpoint)
↓
Route (external access)
OpenShift Builds (optional)
OpenShift provides built-in build capabilities:
- BuildConfigs – Define how to build container images from source
- ImageStreams – Track and manage image versions
- Source-to-Image (S2I) – Build images from source code without writing Dockerfiles
While Builds are available, many organizations prefer using Azure Container Registry for image storage and external CI/CD pipelines for builds.
Container Images and Registries
Container image workflow
Source Code
↓
Build Container Image
↓
Azure Container Registry (ACR)
↓
ARO cluster (pulls image)
↓
Pod runs
ACR integration
ARO integrates with Azure Container Registry for private container image storage. In private clusters, the Microsoft.ContainerRegistry service endpoint on subnets provides direct VNet-integrated access to ACR.
Authentication: Use managed identities or service principals for ACR authentication. For clusters created with managed identities, Workload Identity is available for platform operators.
Image versioning
Recommended: Use immutable image tags for production:
- Semantic versioning:
myapp:v1.2.3 - Git commit SHA:
myapp:a1b2c3d - Date-based:
myapp:2024-01-15
Avoid: Using latest in production. The latest tag is mutable and provides no versioning guarantees.
Image security
- Scan images before deployment
- Use trusted registries (ACR recommended)
- Implement image signing where required
- Regularly update base images
OpenShift Routes and Ingress
Routes vs. Ingress
OpenShift Routes are OpenShift's native ingress solution:
| Aspect | OpenShift Route | Kubernetes Ingress |
|---|---|---|
| Implementation | OpenShift-native | Kubernetes-native |
| TLS termination | Built-in, automatic | Requires configuration |
| Wildcard domains | Supported | Limited |
| Traffic splitting | Supported (weighted) | Limited |
| Security | SCC integration | Standard Kubernetes |
Route configuration
Routes expose services externally with:
- Hostname – Custom or generated domain
- Path – URL path routing
- TLS – Edge, passthrough, or re-encrypt termination
- Weight – Traffic splitting between multiple services
Public vs. private routes
| Route Type | Visibility | Use Case |
|---|---|---|
| Public | Internet-accessible | User-facing applications |
| Private | Cluster-internal only | Internal services, APIs |
| Private with internal ingress | VNet-accessible | Enterprise internal applications |
Custom domains
By default, routes use the ARO-provided domain (*.apps.<random>.<location>.aroapp.io). Custom domains can be configured by:
- Creating a DNS record pointing to the cluster's ingress IP
- Configuring the route with the custom hostname
- Managing TLS certificates (custom or Let's Encrypt)
Storage
Persistent storage in ARO
ARO supports persistent storage for stateful applications through Kubernetes PersistentVolumes (PVs) and PersistentVolumeClaims (PVCs).
Storage classes: ARO provides storage classes for Azure managed disks and Azure Files.
PV limits: The upper limit for the number of PersistentVolumes that can be connected per node is 16.
Storage options
| Storage Type | Access Mode | Use Cases |
|---|---|---|
| Azure Managed Disks | ReadWriteOnce | Databases, stateful applications |
| Azure Files | ReadWriteMany | Shared configuration, content |
| Azure NetApp Files | ReadWriteMany | High-performance enterprise storage |
Stateful applications
For stateful applications:
- Use StatefulSets for ordered deployment and stable network identity
- Use PVCs for persistent data
- Use storage classes that support volume expansion for production
Backup and recovery
OpenShift API for Data Protection (OADP) provides comprehensive backup and recovery for ARO clusters:
- Covers OpenShift applications
- Application-related cluster resources
- Persistent volumes
Important: For production environments, storage classes must have allowVolumeExpansion enabled to allow persistent volumes to be expanded if necessary.
Data externalization
Best practice: Not every workload should store data inside the Kubernetes cluster. Consider:
- Azure SQL or Cosmos DB for relational data
- Azure Storage for blobs and files
- Azure Cache for Redis for caching
- Azure Database services for managed databases
Scaling
Pod scaling (Horizontal Pod Autoscaler)
HPA automatically scales the number of pods based on:
- CPU utilization
- Memory utilization
- Custom metrics (queue depth, request rate, etc.)
Node scaling
ARO supports scaling worker nodes:
- Manual scaling – Add or remove nodes through the Azure portal or CLI
- Cluster Autoscaler – Automatically adjusts node count based on pending pods
Scaling considerations
Resource requests and limits – Set appropriate requests and limits for pods to enable effective scheduling and autoscaling.
Downstream capacity – Scaling pods doesn't solve bottlenecks in databases, APIs, or other downstream services.
Cluster capacity – Ensure the cluster has sufficient capacity for scaling. The cluster requires at least 44 vCPUs for initial creation.
High Availability and Reliability
Availability zones
Deploy ARO clusters across availability zones for zone-level failure protection. This provides resilience if an entire availability zone experiences an outage.
Worker node redundancy
- Use at least three worker nodes for production workloads
- Distribute nodes across availability zones
- Use pod anti-affinity to spread pods across nodes
Control plane availability
The ARO control plane is fully managed with built-in high availability. The joint Microsoft/Red Hat SRE team ensures:
- Redundant control plane components
- Automated failover
- Regular health monitoring
Pod-level reliability
- Use multiple replicas for critical workloads
- Configure pod anti-affinity to avoid single points of failure
- Use Pod Disruption Budgets (PDBs) to maintain availability during voluntary disruptions
- Implement health probes (liveness, readiness, startup)
Application design
- Design applications to be stateless where possible
- Externalize persistent state to Azure services
- Handle failures gracefully with retries and circuit breakers
Disaster Recovery
Backup strategy
Cluster configuration: Back up OpenShift cluster configuration using OADP.
Application manifests: Store application YAML definitions in version control (Git).
Container images: Store images in Azure Container Registry; ensure ACR is backed up or geo-replicated.
Persistent data: Back up persistent volumes and databases regularly.
Configuration: Back up ConfigMaps, Secrets, and other configuration resources.
Multi-region strategy
For critical workloads, deploy clusters across multiple Azure regions:
- Use Azure Front Door or Traffic Manager for global traffic routing
- Use geo-replicated ACR for image availability
- Use geo-replicated databases for data availability
Recovery principle
Do not treat the Kubernetes cluster itself as the only source of truth. Keep application definitions, infrastructure definitions, container images, data, and configuration in recoverable systems (Git, ACR, Azure Storage, databases).
Observability
OpenShift monitoring stack
OpenShift includes a built-in monitoring stack based on Prometheus, Alertmanager, and Grafana. The cluster monitoring stack is managed by the platform and provides:
- Node and pod metrics
- Cluster health indicators
- Alerting capabilities
Note: The ARO operator (aro-operator-master) reverts changes to the monitoring configuration back to a supported configuration, including any changes to retention and storage configuration.
Cluster Observability Operator (COO)
The Cluster Observability Operator (COO) is an optional OpenShift operator that enables administrators to create standalone monitoring stacks that are independently configurable. COO is ideal for users who need high customizability, scalability, and long-term data retention, especially in complex, multi-tenant enterprise environments.
COO can be used to:
- Set up a highly available Prometheus instance that persists metrics
- Enable remote writing of metrics to an Azure Monitor Prometheus workspace
- Create independent monitoring stacks for different services and users
Azure Monitor integration
Log forwarding: Starting from OpenShift Logging version 5.9, OpenShift supports native forwarding to Azure Monitor and Azure Log Analytics, available on clusters running OpenShift 4.13 or higher. This allows you to view and query the logs the platform and your workloads generate in Azure Monitor.
Metric remote write: COO enables remote writing of metrics to Azure Monitor Prometheus workspaces, allowing centralization of metrics in Azure.
The three pillars of observability
Metrics: CPU, memory, network, and custom application metrics Logs: Application logs, container logs, and audit logs Traces: Distributed tracing for microservices
What to monitor
| Role | What to Monitor |
|---|---|
| Developers | Application errors, latency, request rate, pod restarts |
| Platform engineers | Node utilization, cluster capacity, scheduling failures, API health |
| Architects | Availability, dependency health, capacity, failure domains, cost |
CI/CD and DevOps
Production deployment pipeline
Developer
↓
Git (source control)
↓
CI Pipeline (GitHub Actions / Azure DevOps)
↓
Build container image
↓
Security scan
↓
Push to Azure Container Registry
↓
CD Pipeline (GitHub Actions / Azure DevOps / OpenShift Pipelines)
↓
Deploy to OpenShift (oc apply / Helm / GitOps)
↓
Monitoring
↓
Rollback if required
OpenShift Pipelines
OpenShift Pipelines is a Kubernetes-native CI/CD framework based on Tekton. It provides:
- Declarative pipeline definitions as Kubernetes resources
- Reusable tasks and steps
- Integration with OpenShift's developer experience
- CI/CD workflows that run inside the cluster
OpenShift GitOps
OpenShift GitOps is based on ArgoCD and provides GitOps-style continuous delivery. It offers:
- Declarative application delivery
- Git as the single source of truth
- Automatic sync between Git and cluster state
- Least privileged access and version control
CI/CD integration options
| Tool | Integration | Best For |
|---|---|---|
| GitHub Actions | Direct, via OIDC or service principal | GitHub-based workflows |
| Azure DevOps | Native via service connection | Azure-centric teams |
| OpenShift Pipelines | Native, Kubernetes-based | Teams using OpenShift ecosystem |
| OpenShift GitOps | Native, ArgoCD-based | GitOps-driven teams |
Infrastructure as Code
- Use Terraform for ARO cluster provisioning
- Use Bicep or ARM templates for Azure infrastructure
- Use Helm for application packaging and deployment
- Store all IaC in version control
Operators
The Operator pattern
Operators are a Kubernetes-native way to automate the lifecycle of complex applications. The Operator pattern uses Custom Resources to define the desired state and controllers to reconcile the actual state with the desired state.
Custom Resource (desired state)
↓
Operator (controller)
↓
Kubernetes Resources (actual state)
↓
Running Application
Why Operators matter in OpenShift
Operators are central to OpenShift's platform model. They provide:
- Lifecycle automation – Install, upgrade, and manage applications
- Day 2 operations – Backup, recovery, scaling, and configuration
- Vendor integration – Database operators, middleware operators, platform components
Common Operators
- Database operators – PostgreSQL, MySQL, MongoDB, etc.
- Middleware operators – Kafka, RabbitMQ, etc.
- Platform operators – Ingress, monitoring, logging
- Application operators – Custom operators for specific applications
When Operators add value
Operators improve operational consistency when:
- Deploying complex distributed systems
- Managing stateful applications
- Automating day-2 operations
- Standardizing application lifecycle across environments
When Operators add complexity
Operators can introduce additional complexity when:
- Simple applications don't need them
- The Operator itself requires significant management
- Custom Operators are poorly maintained
OpenShift Developer Experience
OpenShift Web Console
The OpenShift web console provides a comprehensive interface for both administrators and developers:
- Developer perspective – Create and deploy applications, view logs, manage routes
- Administrator perspective – Manage cluster resources, projects, RBAC, and operators
OpenShift CLI (oc)
oc extends kubectl with OpenShift-specific commands:
oc new-app– Create applications from source, images, or templatesoc new-project– Create projects with resource quotasoc expose– Create routes to expose servicesoc login– Authenticate to the clusteroc whoami– Show current user
Developer workflow
- Login –
oc loginwith credentials - Create or select project –
oc new-projectoroc project - Deploy application –
oc new-apporoc apply -fwith YAML - Expose application –
oc expose serviceto create a route - Monitor –
oc logs,oc status, or use the web console - Update –
oc apply -fwith updated YAML oroc patch
Focus areas
Developers focus on:
- Application code
- Container images
- Configuration (ConfigMaps, Secrets)
- Deployment manifests
Platform teams provide:
- Cluster infrastructure
- Security and RBAC
- Networking and ingress
- Shared services
ARO for Microservices
ARO is well-suited for enterprise microservices architectures.
Microservices architecture example
External Client
↓
OpenShift Route (public ingress)
↓
API Gateway Service
↓
┌───────────────┬────────────────┬────────────────┐
↓ ↓ ↓ ↓
Orders Payments Users Notifications
Service Service Service Service
↓ ↓ ↓ ↓
└───────────────┴────────────────┴────────────────┘
│
▼
Service Bus (async events)
│
▼
Background Workers
Service boundaries
Each microservice is deployed as a separate OpenShift Deployment:
- Independent scaling based on workload
- Independent deployment with rolling updates
- Independent failure isolation
- Own data store (database per service)
Communication patterns
| Pattern | Implementation | Use Case |
|---|---|---|
| Synchronous HTTP | OpenShift Services (ClusterIP) | Request-response APIs |
| Async messaging | Azure Service Bus / Kafka | Event-driven decoupling |
| Service mesh | OpenShift Service Mesh (Istio) | Advanced traffic management |
Independent deployment and scaling
Each service has its own:
- Deployment with rolling updates
- Horizontal Pod Autoscaler
- Service and Route configuration
Failure isolation
- Failures are contained within each service
- Circuit breakers prevent cascading failures
- Retries and timeouts protect downstream services
ARO for AI and Cloud-Native AI
ARO is increasingly used for AI workloads, especially with native GPU support.
GPU support in ARO
Azure Red Hat OpenShift now supports NVIDIA H100 and H200 GPU-based Azure Virtual Machine SKUs, enabling customers to run large-scale AI, machine learning, and high-performance computing (HPC) workloads. The NVIDIA H200 GPU SKU (Standard_ND96isr_H200_v5) provides 96 vCPUs.
OpenShift AI on ARO addresses GPU workload challenges with flexible, hourly GPU worker node availability, automated scale-out and scale-down capabilities, and a hybrid cloud strategy.
AI use cases on ARO
- AI inference APIs – Deploy trained models as scalable services
- RAG applications – Host Retrieval-Augmented Generation APIs
- AI agents – Deploy agentic AI services
- MLOps platforms – Run ML pipelines and model training
- Data processing – Process large datasets for AI/ML
- Batch inference – Run inference on large batches of data
AI architecture example
Client
↓
OpenShift Route
↓
AI API Service
↓
┌──────────────┬────────────────┬────────────────┐
↓ ↓ ↓ ↓
Azure OpenAI Azure AI Vector Store Data Store
Search
ARO vs. other platforms for AI
| Platform | Best For |
|---|---|
| ARO | Enterprise AI platforms, existing OpenShift skills, complex AI workloads, GPU workloads |
| Azure Container Apps | Simpler AI APIs, event-driven AI, serverless scaling |
| Azure AI services | Managed AI services (OpenAI, Vision, etc.) |
| AKS | Azure-native Kubernetes, custom Kubernetes operators |
When ARO is the right choice
ARO is ideal when:
- Your organization standardizes on OpenShift
- You need enterprise Kubernetes for AI workloads
- You require GPU acceleration with flexible availability
- You have complex AI workloads requiring platform customization
- You're building an enterprise AI platform
ARO and Enterprise Hybrid Cloud
OpenShift consistency
Organizations that use OpenShift on-premises and in other clouds can benefit from ARO's consistency:
- Same OpenShift APIs and platform capabilities
- Same developer experience across environments
- Same operational patterns
- Common tooling and skills
Hybrid cloud architecture
On-Premises OpenShift
↕ (OpenShift APIs)
Azure Red Hat OpenShift
↕ (OpenShift APIs)
OpenShift on Other Clouds
Application portability
OpenShift provides application portability across:
- On-premises OpenShift
- Azure Red Hat OpenShift
- Red Hat OpenShift Service on AWS (ROSA)
- Other OpenShift distributions
The architectural trade-off
| Approach | Benefits | Trade-offs |
|---|---|---|
| Standardize on OpenShift | Consistency, portability, common skills | OpenShift complexity, licensing costs |
| Standardize on Azure-native services | Azure integration, lower cost, simpler | Vendor lock-in, different skill sets |
ARO vs. Azure Kubernetes Service (AKS)
| Dimension | ARO | AKS |
|---|---|---|
| Platform | OpenShift (enterprise Kubernetes) | Vanilla Kubernetes |
| Vendor ecosystem | Red Hat + Microsoft | Microsoft + Kubernetes ecosystem |
| Kubernetes control | Managed | Managed |
| OpenShift capabilities | ✅ Yes (operators, routes, SCCs, console) | ❌ No |
| Developer experience | OpenShift web console + oc CLI | Kubernetes dashboard + kubectl |
| Operators | OpenShift ecosystem (Red Hat, certified, community) | Kubernetes ecosystem |
| Azure integration | Strong | Strong |
| Existing OpenShift skills | Excellent fit | Not applicable |
| Azure-native Kubernetes | Good | Excellent |
| Security defaults | Stricter (SCCs applied by default) | Standard Kubernetes |
| Support | Joint Microsoft/Red Hat | Microsoft |
| Cost | Higher (includes Red Hat licensing) | Lower (free control plane) |
| Enterprise OpenShift standardization | Excellent | Not applicable |
Decision framework
Choose ARO when:
- The organization standardizes on OpenShift
- Existing Red Hat/OpenShift expertise is important
- Hybrid OpenShift consistency matters (on-prem + cloud)
- OpenShift platform capabilities (routes, SCCs, operators) are required
- You need joint Microsoft/Red Hat support
Choose AKS when:
- Azure-native Kubernetes is preferred
- The organization does not need OpenShift
- Kubernetes ecosystem compatibility is sufficient
- Azure integration and Kubernetes flexibility are primary goals
- Cost is a primary concern
ARO vs. Self-Managed OpenShift
| Aspect | Self-Managed OpenShift | ARO |
|---|---|---|
| Control | Full | Limited (platform managed) |
| Operational responsibility | High | Low |
| Infrastructure management | Customer manages | Azure manages |
| Upgrades | Customer plans and executes | Managed by SRE team |
| Availability | Customer responsibility | 99.95% SLA (with availability zones) |
| Networking | Full control | VNet-integrated |
| Security | Customer responsibility | Platform + customer shared |
| Platform maintenance | Customer manages | Automated by platform |
| Cost | Infrastructure + Red Hat subscription | Included in Azure billing |
| Flexibility | Maximum | Limited to ARO capabilities |
Why choose managed
A managed service is preferable for organizations that want OpenShift without owning the full cluster lifecycle. The joint SRE team handles installation, scaling, security, monitoring, and updating of both control plane and worker nodes.
ARO vs. Azure Container Apps
| Aspect | ARO | Azure Container Apps |
|---|---|---|
| Abstraction level | Kubernetes/OpenShift platform | Serverless container platform |
| Kubernetes/OpenShift access | ✅ Full OpenShift API | ❌ No Kubernetes access |
| Developer experience | OpenShift platform | Application-focused |
| Operational complexity | Higher | Lower |
| Scaling | HPA + Cluster Autoscaler | KEDA-based, scale-to-zero |
| Microservices | Full OpenShift support | Built-in service discovery |
| Platform control | High | Limited |
| Networking | Full VNet control | Managed networking |
| Enterprise platform | Yes | Limited |
Decision framework
Choose ARO when:
- You need OpenShift
- You need Kubernetes/OpenShift APIs and platform capabilities
- Enterprise OpenShift standardization matters
- You have advanced platform requirements
- You need GPU support for AI workloads
Choose Container Apps when:
- You want the simplest managed container application platform
- You do not need Kubernetes/OpenShift control
- You are building APIs, workers, or lightweight microservices
- You want scale-to-zero
Cost Considerations
Cost drivers
- Cluster infrastructure – VM costs for worker nodes
- Red Hat licensing – Included in ARO pricing
- Storage – Azure managed disks, Azure Files
- Networking – Data transfer, load balancers
- Monitoring – Log Analytics, Azure Monitor
- Supporting Azure services – ACR, databases, etc.
ARO vs. AKS cost comparison
ARO typically costs approximately 1.8x more than AKS due to the included Red Hat licensing. This premium provides:
- OpenShift platform capabilities
- Joint Microsoft/Red Hat support
- Enterprise-grade security and compliance
- Integrated CI/CD and monitoring
Cost optimization
- Right-size worker nodes based on workload requirements
- Use autoscaling to match capacity with demand
- Use spot instances for non-production workloads (where supported)
- Separate development and production clusters to right-size each
- Monitor and clean up unused resources
Production Best Practices
Cluster design
- Use private clusters for production workloads
- Deploy across availability zones for high availability
- Use separate node pools for system and user workloads
- Add infrastructure nodes for platform workloads
- Right-size worker nodes based on workload characteristics
Identity and security
- Integrate Microsoft Entra ID for authentication
- Use Azure Managed Identities instead of service principals
- Apply least-privilege OpenShift RBAC
- Use Security Context Constraints to enforce pod security
- Enable audit logging
- Use Azure Key Vault for secrets
Networking
- Plan VNet and subnet sizing carefully
- Use private clusters for sensitive workloads
- Configure egress appropriately (NAT Gateway or Azure Firewall)
- Enable service endpoints for ACR access in private clusters
- Use custom domains for production routes
Application management
- Use immutable image versions (avoid
latest) - Scan container images for vulnerabilities
- Use GitOps for declarative application delivery
- Implement health probes (liveness, readiness, startup)
- Use Pod Disruption Budgets for critical workloads
Observability
- Set up log forwarding to Azure Monitor
- Configure metric remote write to Azure Monitor
- Set up alerts for critical conditions
- Monitor cluster and application health
Operations
- Test upgrades in staging before production
- Back up cluster configuration and persistent data
- Document dependencies and failure modes
- Use Infrastructure as Code for cluster provisioning
- Regularly review permissions and access
Common Mistakes
Treating ARO as ordinary Kubernetes
OpenShift adds significant capabilities and constraints beyond vanilla Kubernetes. Understand OpenShift-specific concepts like Projects, Routes, and Security Context Constraints.
Assuming ARO is identical to AKS
ARO and AKS serve different purposes. ARO provides OpenShift with enterprise features; AKS provides vanilla Kubernetes. Choose based on your requirements.
Creating excessive clusters instead of using projects
Projects provide logical isolation within a cluster. Creating separate clusters for every application increases cost and management complexity.
Using cluster-admin permissions for normal development
Cluster-admin provides full access to the entire cluster. Use project-scoped permissions for developers.
Ignoring OpenShift RBAC
Azure RBAC controls Azure resource management. OpenShift RBAC controls access to cluster resources. Both are needed.
Running privileged containers unnecessarily
OpenShift applies restrictive SCCs by default. Don't grant privileged access unless absolutely required.
Storing secrets inside container images
Secrets in images are exposed to anyone with image access. Use OpenShift Secrets or Azure Key Vault.
Using mutable image tags (latest) in production
latest provides no versioning guarantees. Use immutable tags for production.
Ignoring resource requests and limits
Without resource limits, workloads can consume all cluster resources. Always set requests and limits.
Ignoring cluster capacity
Cluster creation requires at least 44 vCPUs. Monitor capacity and plan for growth.
Assuming pod autoscaling solves all capacity problems
HPA scales pods but doesn't add nodes. Use Cluster Autoscaler for node scaling.
Exposing internal services publicly
Use internal routes for services that shouldn't be publicly accessible.
Ignoring network egress
Private clusters need configured egress for internet access. Plan egress routing.
Treating high availability as disaster recovery
HA protects against node/zone failures. DR protects against region failures. Different strategies needed.
Keeping state inside pods
Pods are ephemeral. Use PersistentVolumes or Azure services for persistent state.
Failing to monitor node and pod health
Monitor both cluster and application health. Set up alerts for critical conditions.
Choosing ARO when a simpler managed container service is sufficient
ARO adds complexity and cost. Consider Container Apps for simpler workloads.
Choosing ARO without considering existing organizational OpenShift skills
ARO requires OpenShift knowledge. If the team doesn't have OpenShift experience, AKS or Container Apps may be better.
Underestimating OpenShift operational complexity
ARO reduces platform management but doesn't eliminate it. You still need to manage workloads, networking, security, and operations.
Troubleshooting
Cluster access failure
- Symptoms: Cannot access cluster API or console
- Likely causes: Network connectivity, expired credentials, RBAC issues
- Diagnose: Check network connectivity; verify authentication; check Azure RBAC
- Fix: Use
az aro list-credentialsto retrieve credentials; verify network access
Authentication problems
- Symptoms: Cannot log in; unauthorized errors
- Likely causes: Incorrect identity provider configuration; expired client secret
- Diagnose: Check OpenShift OAuth configuration; verify Entra ID app registration
- Fix: Update OAuth configuration; regenerate client secret
RBAC denial
- Symptoms: "Forbidden" errors when accessing resources
- Likely causes: Missing RBAC permissions
- Diagnose: Check OpenShift RBAC roles and bindings
- Fix: Grant appropriate roles (view, edit, admin) at the project or cluster scope
Pod stuck in Pending
- Symptoms: Pod remains in Pending state
- Likely causes: Insufficient cluster capacity; resource requests exceed available capacity
- Diagnose: Check pod events (
oc describe pod); check node capacity - Fix: Add nodes; reduce resource requests; use different node pool
Pod CrashLoopBackOff
- Symptoms: Pod crashes and restarts repeatedly
- Likely causes: Application errors; missing dependencies; configuration issues
- Diagnose: Check pod logs (
oc logs); check pod events - Fix: Fix application code; correct configuration; increase resource limits
Image pull failure
- Symptoms: Pod shows ImagePullBackOff or ErrImagePull
- Likely causes: Image doesn't exist; incorrect tag; ACR authentication failure
- Diagnose: Check image name and tag; verify ACR access
- Fix: Push the image; correct the tag; ensure pull secret is configured
Route unavailable
- Symptoms: Route doesn't respond to requests
- Likely causes: Route configuration issues; service not running; TLS misconfiguration
- Diagnose: Check route status (
oc get routes); verify service endpoints - Fix: Correct route configuration; ensure service is running; update TLS certificates
Service-to-service communication failure
- Symptoms: Services can't communicate with each other
- Likely causes: Incorrect service name; network policy blocking traffic
- Diagnose: Verify service DNS resolution; check network policies
- Fix: Use correct service names; configure network policies to allow traffic
Storage mount failure
- Symptoms: PVCs not binding; pod can't mount volumes
- Likely causes: Storage class not available; insufficient permissions; quota exceeded
- Diagnose: Check PVC status (
oc get pvc); verify storage class - Fix: Use correct storage class; ensure sufficient quota; check permissions
High CPU or memory usage
- Symptoms: Nodes or pods using excessive resources
- Likely causes: Application issues; insufficient resource limits; node under-provisioned
- Diagnose: Check node and pod metrics; review application behavior
- Fix: Increase resource limits; scale out; fix application performance issues
Autoscaling not behaving as expected
- Symptoms: HPA not scaling; scaling too aggressively or not at all
- Likely causes: Metrics not available; incorrect configuration; insufficient cluster capacity
- Diagnose: Check HPA status (
oc describe hpa); verify metrics-server - Fix: Correct HPA configuration; ensure metrics are available; add cluster capacity
Practical Learning Path
-
Learn containers and Docker – Basic container concepts, Dockerfiles, image building
-
Learn Kubernetes fundamentals – Pods, Deployments, Services, Namespaces, ConfigMaps, Secrets
-
Understand OpenShift concepts – Projects, Routes, Security Context Constraints, Operators
-
Create an ARO cluster – Use Azure CLI or portal; understand prerequisites (vCPUs, subnets)
-
Explore the OpenShift Console – Developer and Administrator perspectives
-
Deploy a containerized application – Using
oc new-appor YAML manifests -
Learn Projects and RBAC – Create projects; assign roles (view, edit, admin)
-
Configure Routes and Services – Expose applications externally with routes
-
Configure persistent storage – PVs and PVCs for stateful applications
-
Integrate Azure Container Registry – Push images; configure pull secrets
-
Integrate Microsoft Entra ID – Configure OIDC identity provider
-
Configure autoscaling – HPA for pods; Cluster Autoscaler for nodes
-
Implement CI/CD – OpenShift Pipelines or GitOps
-
Add monitoring and logging – COO for custom monitoring; log forwarding to Azure Monitor
-
Secure workloads – SCCs, network policies, secrets management
-
Build a microservices architecture – Multiple services, internal communication, routes
-
Explore Operators – Understand Operator pattern; deploy common operators
-
Design hybrid OpenShift architectures – Connect ARO with on-premises OpenShift
-
Evaluate ARO vs. AKS – Understand the decision framework
-
Apply ARO to AI/cloud-native workloads – GPU nodes, model inference, RAG
Interview and Architecture Questions
What is Azure Red Hat OpenShift? Azure Red Hat OpenShift is a fully managed, enterprise-grade Kubernetes platform jointly developed, operated, and supported by Red Hat and Microsoft on the Azure cloud. It provides OpenShift capabilities without the complexity of managing the underlying infrastructure.
How is ARO different from AKS? ARO provides OpenShift with enterprise features (routes, SCCs, operators, web console) and joint Microsoft/Red Hat support. AKS provides vanilla Kubernetes with Azure-native integrations and lower cost. ARO typically costs more due to Red Hat licensing.
What does Microsoft manage in ARO? Microsoft and Red Hat manage the installation, scaling, security, monitoring, and updating of both control plane and worker nodes. The joint SRE team ensures high availability and resilience.
What does the customer manage? Customers manage worker node configuration, scaling, application workloads, networking (VNet), identity integration, storage configuration, and monitoring setup.
What is the difference between Kubernetes and OpenShift? OpenShift builds on Kubernetes with an enterprise-ready platform including container management, automation, networking, CI/CD, monitoring, registry, and authentication—all tested together. Kubernetes is CaaS; OpenShift/ARO is PaaS.
What is an OpenShift Project? A Project is OpenShift's abstraction over Kubernetes namespaces with additional features: resource quotas, limit ranges, project-scoped RBAC, and network policies.
What is an OpenShift Route? A Route is OpenShift's native ingress solution that exposes services externally with hostname-based routing, path-based routing, TLS termination, and traffic splitting.
What is an Operator? An Operator is a Kubernetes-native pattern for automating the lifecycle of complex applications using Custom Resources and controllers. Operators manage installation, upgrades, backup, and recovery.
How does ARO integrate with Azure networking? ARO is deployed into a customer's Azure VNet with dedicated subnets for master and worker nodes. It supports private clusters, service endpoints, and Private Link.
How does identity work in ARO? ARO integrates with Microsoft Entra ID as an OIDC identity provider. Users authenticate via Entra ID; group claims map Entra ID groups to OpenShift RBAC. Managed Identities and Workload Identity are supported for platform operators.
How do Azure RBAC and OpenShift RBAC differ? Azure RBAC controls who can manage ARO resources in Azure. OpenShift RBAC controls who can access cluster resources (projects, pods, services). Both are needed for comprehensive access control.
How would you secure an ARO cluster? Use private clusters, integrate Entra ID, apply least-privilege RBAC, use SCCs, enable audit logging, use network policies, scan images, use Azure Key Vault for secrets, and monitor regularly.
How would you design a production ARO architecture? Use private clusters across availability zones, separate node pools for system and user workloads, add infrastructure nodes, integrate Entra ID, use ACR with private endpoint, implement monitoring with Azure Monitor, and use GitOps for deployment.
When would you choose ARO instead of AKS? Choose ARO when your organization standardizes on OpenShift, needs OpenShift capabilities, wants hybrid OpenShift consistency, or requires joint Microsoft/Red Hat support.
When would you choose Container Apps instead of ARO? Choose Container Apps when you want the simplest managed container platform, don't need Kubernetes/OpenShift control, and are building APIs or lightweight microservices.
How would you design hybrid OpenShift? Use OpenShift APIs for consistency across on-premises OpenShift and ARO. Use the same application definitions, RBAC patterns, and deployment processes. Use Azure networking (VPN, ExpressRoute) for connectivity.
How would you deploy AI workloads on ARO? Use GPU-enabled worker nodes (NVIDIA H100/H200). Deploy model-serving containers. Use OpenShift AI or custom Operators for ML pipelines. Integrate with Azure AI services and vector stores.
How would you troubleshoot a pod stuck in Pending?
Check pod events (oc describe pod). Check node capacity and resource availability. Verify resource requests don't exceed available capacity. Check for node taints or tolerations.
How would you troubleshoot an image pull failure? Check the image name and tag. Verify the image exists in ACR. Check pull secret configuration. For private clusters, verify ACR service endpoint is enabled.
Key Takeaways
Azure Red Hat OpenShift (ARO) is a fully managed, enterprise-grade Kubernetes platform jointly developed, operated, and supported by Red Hat and Microsoft on Azure.
ARO provides OpenShift's enterprise capabilities on Azure – routes, SCCs, operators, built-in CI/CD, monitoring, and a developer-friendly web console.
The platform is managed by a joint SRE team from Red Hat and Microsoft, handling installation, scaling, security, monitoring, and updates. This significantly reduces operational burden.
ARO is built on Azure infrastructure – VMs, network security groups, and storage accounts are deployed into your Azure subscription. You integrate with Azure VNets, subnets, and Azure services.
Private clusters require careful networking configuration, including service endpoints for ACR access and custom egress routing via NAT Gateway or Azure Firewall.
Identity integrates with Microsoft Entra ID as an OIDC provider. Managed Identities and Workload Identity provide secret-free authentication for platform operators.
OpenShift adds Security Context Constraints (SCCs) that enforce stricter security defaults than vanilla Kubernetes. Pods require explicit opt-in for privileged operations.
ARO supports NVIDIA H100 and H200 GPUs for large-scale AI, machine learning, and HPC workloads. OpenShift AI provides flexible GPU worker node availability.
Observability integrates with Azure Monitor – log forwarding is native in OpenShift 4.13+, and the Cluster Observability Operator enables remote metric writes.
Choose ARO when your organization standardizes on OpenShift, needs OpenShift capabilities, or requires hybrid OpenShift consistency. Choose AKS for Azure-native Kubernetes without OpenShift.
Production best practices include: use private clusters, deploy across availability zones, integrate Entra ID, apply least-privilege RBAC, use immutable image versions, implement monitoring, and use GitOps for declarative deployment.