Skip to main content

Azure Red Hat OpenShift (ARO) Guide: Architecture, Use Cases, and Best Practices

Kubernetes has become the de facto standard for container orchestration. It provides powerful capabilities for scheduling, scaling, and managing containerized workloads. But running Kubernetes in production—especially at enterprise scale—is operationally complex. Managing the control plane, configuring networking, securing the cluster, handling upgrades, and integrating with existing identity and governance systems requires specialized expertise that many organizations struggle to maintain.

This is where enterprise Kubernetes platforms like OpenShift add value. OpenShift builds on Kubernetes with additional capabilities: integrated CI/CD pipelines, built-in monitoring and logging, a developer-friendly web console, security context constraints, operators, and a curated ecosystem of tools and services. It provides a more complete platform experience out of the box.

Azure Red Hat OpenShift (ARO) brings this enterprise platform to Azure as a fully managed service. ARO is a fully managed, enterprise-grade Kubernetes platform jointly developed, operated, and supported by Red Hat and Microsoft. It allows you to deploy fully managed OpenShift clusters without the complexity of building, maintaining, or securing the underlying infrastructure.

The joint engineering and support model is a key differentiator: a dedicated Site Reliability Engineering (SRE) team from Red Hat and Microsoft works together to ensure high availability and resilience of your clusters. This means you get a seamless experience with integrated support from both vendors.

ARO is different from:

  • Self-managed OpenShift on Azure VMs – You manage the entire OpenShift lifecycle yourself
  • Azure Kubernetes Service (AKS) – Vanilla Kubernetes without the OpenShift platform layer
  • Azure Container Apps – A higher-level abstraction for containerized applications without Kubernetes control
  • Azure Red Hat Enterprise Linux – A Linux distribution, not a container platform

ARO provides an enterprise Kubernetes platform that combines Red Hat's OpenShift capabilities with Azure's infrastructure and services, while significantly reducing the operational burden of managing the underlying platform.

What Is Azure Red Hat OpenShift?​

Azure Red Hat OpenShift (ARO) is a fully managed OpenShift service running on Azure that provides an enterprise Kubernetes platform while reducing the operational burden of managing the underlying OpenShift control plane and infrastructure.

The conceptual stack​

Application Layer
↓
OpenShift Platform Layer (Developer tools, Operators, CI/CD, Monitoring)
↓
Kubernetes Layer (Container orchestration)
↓
Azure Infrastructure Layer (VMs, networking, storage)
↓
Azure Physical Infrastructure

What OpenShift adds to Kubernetes​

OpenShift is not merely "Kubernetes with a UI." It adds several enterprise-grade capabilities on top of Kubernetes:

  • OpenShift Web Console – A comprehensive web-based interface for both administrators and developers
  • OpenShift CLI (oc) – An enhanced CLI that extends kubectl with OpenShift-specific commands
  • Projects – OpenShift's abstraction over Kubernetes namespaces with additional security and resource controls
  • Routes – OpenShift's native ingress solution for exposing services externally
  • Security Context Constraints (SCCs) – Fine-grained control over pod security permissions
  • Operators – A framework for automating the lifecycle of complex applications
  • OpenShift GitOps and Pipelines – Built-in CI/CD capabilities based on ArgoCD and Tekton
  • Integrated container registry – A built-in registry for storing container images
  • Integrated monitoring – Prometheus-based monitoring stack with Thanos for long-term storage

ARO vs. Kubernetes​

While Kubernetes is considered CaaS (Container as a Service), OpenShift and ARO fall under the category PaaS (Platform as a Service). Unlike basic Kubernetes, OpenShift includes pre-integrated components: container management, automation, networking, CI/CD, monitoring, registry, and authentication—all tested together.

ARO vs. self-managed OpenShift​

OpenShift offers two primary deployment models:

AspectSelf-Managed OpenShiftAzure Red Hat OpenShift (ARO)
ManagementOrganization manages installation, updates, and managementFully managed by Red Hat and Microsoft
InfrastructureOrganization manages underlying infrastructureAzure infrastructure managed by Microsoft
SupportRed Hat supportJoint Microsoft/Red Hat support
BillingDirect Red Hat subscription + infrastructureBilled through Azure subscription
DeploymentOn any supported infrastructureExclusive to Azure

Azure Red Hat OpenShift at a Glance​

CapabilityPurpose
Managed OpenShiftRun enterprise OpenShift without managing the entire platform yourself
KubernetesContainer orchestration foundation
OpenShift ConsoleWeb-based administration and developer experience
ProjectsNamespace-oriented application isolation with enhanced security
RoutesExpose applications externally with built-in TLS
OperatorsManage platform and application components
Azure networkingIntegrate clusters with Azure VNets and subnets
Microsoft Entra IDIdentity integration for authentication
Azure MonitorMonitoring and observability integration via Cluster Logging Forwarder and Cluster Observability Operator
Azure Container RegistryPrivate container image storage and management
Azure StoragePersistent application data via PVs and PVCs
Azure Key VaultSecrets and security integration
AutoscalingScale workloads with Horizontal Pod Autoscaler
AvailabilitySupport resilient production cluster architectures with availability zones
OpenShift APIs and CLIAutomation and developer workflows
GPU SupportSupport for NVIDIA H100 and H200 GPU-based Azure VM SKUs for AI/ML workloads

ARO Architecture​

Major layers​

Azure Subscription
↓
Resource Group
↓
Virtual Network (with subnets)
↓
ARO Cluster
├── Control Plane (master nodes, managed by Azure/Red Hat)
├── Infrastructure Nodes (optional, customer-configurable)
└── Worker Nodes (customer-managed)
↓
OpenShift Workloads
├── Projects / Namespaces
├── Deployments
├── Pods
├── Services
├── Routes
└── Operators

Core components​

ARO is built on Azure infrastructure services, including virtual machines, network security groups, and storage accounts, all deployed directly into your Azure subscription.

Operating System: ARO runs on Red Hat Enterprise Linux CoreOS (RHCOS), providing a secure, immutable OS optimized for running containers. RHCOS is designed specifically for container workloads with automated updates and minimal attack surface.

Control plane: The OpenShift control plane includes the API server, etcd (the cluster's key-value store), scheduler, and controller manager. In ARO, the control plane is fully managed by the joint Microsoft/Red Hat SRE team. You don't have direct access to control plane nodes or etcd.

Infrastructure nodes: ARO supports adding infrastructure nodes to host platform workloads like Ingress controllers, the registry, and cluster monitoring. This helps with larger clusters that have resource contention between user workloads and infrastructure workloads such as Prometheus.

Worker nodes: Worker nodes run your application workloads. You choose the VM size, number of nodes, and scaling behavior.

Responsibility model​

ComponentManaged By
OpenShift control planeMicrosoft + Red Hat
Worker nodes (VMs)Customer (configuration, scaling)
Node OS (RHCOS)Automated updates by platform
Cluster networkingCustomer (VNet configuration)
OpenShift API and consoleMicrosoft + Red Hat
Application workloadsCustomer
Storage configurationCustomer
Identity integrationCustomer
Monitoring configurationCustomer

ARO Cluster Design​

Key design decisions​

Region – Choose Azure regions that support ARO. Cluster creation requires at least 44 vCPUs to create and run an OpenShift cluster.

Availability zones – Deploy ARO clusters across availability zones where supported to increase resilience.

Virtual network and subnets – ARO requires a virtual network with two empty subnets: one for master (control plane) nodes and one for worker nodes. You can create a new VNet or use an existing one. Each ARO cluster should use separate or dedicated subnets to avoid potential conflicts.

Public vs. private clusters:

AspectPublic ClusterPrivate Cluster
API server visibilityPublic endpointPrivate endpoint only
Ingress visibilityPublicPrivate
EgressDefault: LoadBalancer with internet egressRestricted; requires VNet-integrated paths
Best forSandbox, developmentProduction, regulated workloads

Opt for a public cluster only in situations like a "sandbox cluster" or where establishing a private method for console and API access is not feasible or desired.

Private cluster considerations: In a private ARO cluster, the control plane nodes don't have public internet access. The Microsoft.ContainerRegistry service endpoint on the master subnet is required so that private control-plane nodes can reach Azure Container Registry over the virtual network without using public IP connectivity. For private clusters, configuring the service endpoint on the master subnet is a prerequisite so the control plane can reliably access ACR without relying on public internet egress.

Multi-cluster enterprise strategy​

For enterprise environments, consider:

  • Development clusters – Smaller, public (or limited private access), lower cost
  • Test clusters – Medium size, representative of production
  • Production clusters – Private, multi-zone, high availability
  • Multi-cluster management – Use OpenShift GitOps or fleet management for consistency

When to use separate clusters vs. projects​

FactorSeparate ClustersSeparate Projects
IsolationStrongest (network, security, failure)Namespace-level isolation
CostHigher (multiple clusters)Lower (shared cluster)
ManagementMore complexSimpler
ComplianceDifferent compliance requirementsSame compliance boundary
Team autonomyFull cluster controlProject-level control

Networking​

Virtual network requirements​

ARO requires a virtual network with two empty subnets:

  • Master subnet – For control plane nodes (minimum /23 recommended)
  • Worker subnet – For worker nodes (minimum /23 recommended)

The subnets must be empty when creating the cluster; they cannot contain existing resources.

Egress configuration​

By default, public and private clusters have --outbound-type defined to LoadBalancer, meaning all clusters have open egress to the internet through the public load balancer.

To change the default behavior and restrict internet egress, set --outbound-type to UserDefinedRouting during cluster creation and set up traffic to run through a firewall solution, such as Azure Firewall or Azure NAT Gateway.

Egress options:

  • NAT Gateway – Replaces routes to go through Azure NAT Gateway for egress instead of the LoadBalancer
  • Azure Firewall – Routes egress traffic through Azure Firewall with granular rule control

Private clusters and ACR access​

For private clusters, the Microsoft.ContainerRegistry service endpoint on the master and worker subnets provides a direct, VNet-integrated path to ACR. This is required because private control-plane and worker nodes don't have public internet access and need to pull container images for platform components and workloads.

Service endpoints and VNet encryption​

ARO version 4.18+ supports installing clusters with virtual network encryption. In this version, the dependency on service endpoints has been removed, and new clusters won't create service endpoints on the VNet.

DNS and custom domains​

By default, ARO uses self-signed certificates for all routes created on *.apps.<random>.<location>.aroapp.io. Many organizations want to use their own custom domains for applications. ARO supports custom domain configuration for routes.

Network security​

  • Use Network Security Groups (NSGs) to control traffic at the subnet level
  • Use Azure Private Link to access ARO cluster API endpoints and Kubernetes LoadBalancer-type services
  • Configure private endpoints for supported Azure resources to establish private access points
  • ARO service supports deployment to customer virtual networks

Identity and Access Management​

Identity architecture​

ARO supports multiple identity providers, with Microsoft Entra ID (formerly Azure Active Directory) being the primary recommended approach.

Microsoft Entra ID
↓
OpenShift Authentication (OAuth)
↓
OpenShift RBAC
↓
Project
↓
Application

Microsoft Entra ID integration​

ARO can be configured to use Microsoft Entra ID as an OpenID Connect (OIDC) identity provider. The integration involves:

  1. Registering an application in Azure AD for authentication
  2. Configuring the application registration to include optional claims (email, preferred_username) and group claims in tokens
  3. Configuring the ARO cluster to use Azure AD as the identity provider
  4. Granting permissions to individual users or groups

Group claims: OpenShift 4.10+ supports OpenID Connect group claim functionality, allowing an identity provider to provide a user's group membership for use within OpenShift. This enables mapping Azure AD security groups to OpenShift roles.

Azure RBAC vs. OpenShift RBAC vs. Kubernetes RBAC​

AspectAzure RBACOpenShift RBACKubernetes RBAC
ScopeAzure resource managementOpenShift cluster resourcesKubernetes API resources
PurposeControl who can manage ARO resourcesControl access to OpenShift resourcesControl access to Kubernetes resources
UsersAzure AD users and service principalsOpenShift users and service accountsService accounts
IntegrationAzure-nativeEntra ID + OpenShiftOpenShift-native

Azure RBAC alone does not replace OpenShift RBAC. You need both: Azure RBAC controls who can create and manage ARO clusters, while OpenShift RBAC controls who can deploy and manage applications within the cluster.

Managed Identities and Workload Identity​

ARO supports deploying clusters using Azure Managed Identities instead of service principals, enabling Workload Identity for platform operators.

Benefits of Managed Identity with Workload Identity:

  • Enhanced Security: No service principal secrets to manage; no service principal secrets stored in the cluster
  • Streamlined Operations: Automatic credential rotation via Azure Managed Identity; no manual secret rotation required
  • Workload Identity: Platform operators use federated credentials with Azure
  • Compliance: Meets security requirements for secret-free authentication; aligns with Azure security best practices

Microsoft Entra Workload ID is available in OpenShift clusters configured to use short-term credentials, starting with ARO version 4.16 and later when originally created with managed identities.

Least privilege​

Apply least-privilege principles:

  • Use separate service accounts for different workloads
  • Grant only the permissions required for each application
  • Regularly review and audit permissions
  • Use OpenShift RBAC to limit access to projects and resources

Security Architecture​

Security by default​

ARO enforces security best practices by default, with automated updates, integrated monitoring, and compliance controls, making it suitable for running sensitive or regulated workloads. ARO provides stricter security defaults, including Security Context Constraints (SCCs).

Security Context Constraints (SCCs)​

SCCs are OpenShift's mechanism for controlling pod-level security permissions. They govern:

  • Whether pods can run as root
  • Which capabilities are available
  • Access to host resources
  • Filesystem permissions

Unlike vanilla Kubernetes, OpenShift applies restrictive SCCs by default, requiring explicit opt-in for privileged operations.

Network policies​

Use OpenShift network policies to control traffic between pods and services. Network policies provide:

  • Pod-to-pod communication control
  • Namespace-level isolation
  • Service segmentation

Container image security​

  • Scan images for vulnerabilities before deployment
  • Use trusted registries (ACR recommended)
  • Use image signing where applicable
  • Avoid latest tags in production

Secrets management​

  • Use OpenShift Secrets for sensitive configuration
  • Integrate with Azure Key Vault for enterprise secret management
  • Never store secrets in container images or source code
  • Rotate secrets regularly

Audit logging​

  • Enable audit logging for both Azure and OpenShift
  • Monitor for suspicious activity
  • Use Azure Monitor for centralized log collection

OpenShift Projects and Namespaces​

Projects as OpenShift namespaces​

OpenShift Projects are an abstraction over Kubernetes namespaces with additional features:

  • Resource quotas – Limit CPU, memory, and storage usage
  • Limit ranges – Set default resource requests and limits
  • RBAC – Project-scoped roles and bindings
  • Network policies – Per-project network isolation

Organizing with projects​

Projects can be used to organize:

  • Teams – Each team gets its own project(s)
  • Applications – Each application has a dedicated project
  • Environments – Development, test, and production projects
  • Tenants – Multi-tenant SaaS applications

Resource quotas​

Resource quotas prevent any single project from consuming all cluster resources:

  • CPU and memory limits
  • Persistent volume claim limits
  • Object count limits (pods, services, etc.)

When projects are not enough​

Projects provide logical isolation but not physical isolation. For stronger isolation:

  • Use separate node pools with taints and tolerations
  • Use separate clusters for different compliance requirements
  • Use network policies for traffic segmentation

Applications and Workloads​

Workload types​

ARO supports all standard Kubernetes workload types:

  • Deployments – Stateless applications with rolling updates
  • StatefulSets – Stateful applications with persistent storage
  • Jobs – One-off batch processing
  • CronJobs – Scheduled batch processing
  • DaemonSets – Node-level agents and daemons

Application lifecycle​

Container Image (from ACR or other registry)
↓
Deployment (defines desired state)
↓
ReplicaSet (maintains replica count)
↓
Pod (running container instance)
↓
Service (stable network endpoint)
↓
Route (external access)

OpenShift Builds (optional)​

OpenShift provides built-in build capabilities:

  • BuildConfigs – Define how to build container images from source
  • ImageStreams – Track and manage image versions
  • Source-to-Image (S2I) – Build images from source code without writing Dockerfiles

While Builds are available, many organizations prefer using Azure Container Registry for image storage and external CI/CD pipelines for builds.

Container Images and Registries​

Container image workflow​

Source Code
↓
Build Container Image
↓
Azure Container Registry (ACR)
↓
ARO cluster (pulls image)
↓
Pod runs

ACR integration​

ARO integrates with Azure Container Registry for private container image storage. In private clusters, the Microsoft.ContainerRegistry service endpoint on subnets provides direct VNet-integrated access to ACR.

Authentication: Use managed identities or service principals for ACR authentication. For clusters created with managed identities, Workload Identity is available for platform operators.

Image versioning​

Recommended: Use immutable image tags for production:

  • Semantic versioning: myapp:v1.2.3
  • Git commit SHA: myapp:a1b2c3d
  • Date-based: myapp:2024-01-15

Avoid: Using latest in production. The latest tag is mutable and provides no versioning guarantees.

Image security​

  • Scan images before deployment
  • Use trusted registries (ACR recommended)
  • Implement image signing where required
  • Regularly update base images

OpenShift Routes and Ingress​

Routes vs. Ingress​

OpenShift Routes are OpenShift's native ingress solution:

AspectOpenShift RouteKubernetes Ingress
ImplementationOpenShift-nativeKubernetes-native
TLS terminationBuilt-in, automaticRequires configuration
Wildcard domainsSupportedLimited
Traffic splittingSupported (weighted)Limited
SecuritySCC integrationStandard Kubernetes

Route configuration​

Routes expose services externally with:

  • Hostname – Custom or generated domain
  • Path – URL path routing
  • TLS – Edge, passthrough, or re-encrypt termination
  • Weight – Traffic splitting between multiple services

Public vs. private routes​

Route TypeVisibilityUse Case
PublicInternet-accessibleUser-facing applications
PrivateCluster-internal onlyInternal services, APIs
Private with internal ingressVNet-accessibleEnterprise internal applications

Custom domains​

By default, routes use the ARO-provided domain (*.apps.<random>.<location>.aroapp.io). Custom domains can be configured by:

  1. Creating a DNS record pointing to the cluster's ingress IP
  2. Configuring the route with the custom hostname
  3. Managing TLS certificates (custom or Let's Encrypt)

Storage​

Persistent storage in ARO​

ARO supports persistent storage for stateful applications through Kubernetes PersistentVolumes (PVs) and PersistentVolumeClaims (PVCs).

Storage classes: ARO provides storage classes for Azure managed disks and Azure Files.

PV limits: The upper limit for the number of PersistentVolumes that can be connected per node is 16.

Storage options​

Storage TypeAccess ModeUse Cases
Azure Managed DisksReadWriteOnceDatabases, stateful applications
Azure FilesReadWriteManyShared configuration, content
Azure NetApp FilesReadWriteManyHigh-performance enterprise storage

Stateful applications​

For stateful applications:

  • Use StatefulSets for ordered deployment and stable network identity
  • Use PVCs for persistent data
  • Use storage classes that support volume expansion for production

Backup and recovery​

OpenShift API for Data Protection (OADP) provides comprehensive backup and recovery for ARO clusters:

  • Covers OpenShift applications
  • Application-related cluster resources
  • Persistent volumes

Important: For production environments, storage classes must have allowVolumeExpansion enabled to allow persistent volumes to be expanded if necessary.

Data externalization​

Best practice: Not every workload should store data inside the Kubernetes cluster. Consider:

  • Azure SQL or Cosmos DB for relational data
  • Azure Storage for blobs and files
  • Azure Cache for Redis for caching
  • Azure Database services for managed databases

Scaling​

Pod scaling (Horizontal Pod Autoscaler)​

HPA automatically scales the number of pods based on:

  • CPU utilization
  • Memory utilization
  • Custom metrics (queue depth, request rate, etc.)

Node scaling​

ARO supports scaling worker nodes:

  • Manual scaling – Add or remove nodes through the Azure portal or CLI
  • Cluster Autoscaler – Automatically adjusts node count based on pending pods

Scaling considerations​

Resource requests and limits – Set appropriate requests and limits for pods to enable effective scheduling and autoscaling.

Downstream capacity – Scaling pods doesn't solve bottlenecks in databases, APIs, or other downstream services.

Cluster capacity – Ensure the cluster has sufficient capacity for scaling. The cluster requires at least 44 vCPUs for initial creation.

High Availability and Reliability​

Availability zones​

Deploy ARO clusters across availability zones for zone-level failure protection. This provides resilience if an entire availability zone experiences an outage.

Worker node redundancy​

  • Use at least three worker nodes for production workloads
  • Distribute nodes across availability zones
  • Use pod anti-affinity to spread pods across nodes

Control plane availability​

The ARO control plane is fully managed with built-in high availability. The joint Microsoft/Red Hat SRE team ensures:

  • Redundant control plane components
  • Automated failover
  • Regular health monitoring

Pod-level reliability​

  • Use multiple replicas for critical workloads
  • Configure pod anti-affinity to avoid single points of failure
  • Use Pod Disruption Budgets (PDBs) to maintain availability during voluntary disruptions
  • Implement health probes (liveness, readiness, startup)

Application design​

  • Design applications to be stateless where possible
  • Externalize persistent state to Azure services
  • Handle failures gracefully with retries and circuit breakers

Disaster Recovery​

Backup strategy​

Cluster configuration: Back up OpenShift cluster configuration using OADP.

Application manifests: Store application YAML definitions in version control (Git).

Container images: Store images in Azure Container Registry; ensure ACR is backed up or geo-replicated.

Persistent data: Back up persistent volumes and databases regularly.

Configuration: Back up ConfigMaps, Secrets, and other configuration resources.

Multi-region strategy​

For critical workloads, deploy clusters across multiple Azure regions:

  • Use Azure Front Door or Traffic Manager for global traffic routing
  • Use geo-replicated ACR for image availability
  • Use geo-replicated databases for data availability

Recovery principle​

Do not treat the Kubernetes cluster itself as the only source of truth. Keep application definitions, infrastructure definitions, container images, data, and configuration in recoverable systems (Git, ACR, Azure Storage, databases).

Observability​

OpenShift monitoring stack​

OpenShift includes a built-in monitoring stack based on Prometheus, Alertmanager, and Grafana. The cluster monitoring stack is managed by the platform and provides:

  • Node and pod metrics
  • Cluster health indicators
  • Alerting capabilities

Note: The ARO operator (aro-operator-master) reverts changes to the monitoring configuration back to a supported configuration, including any changes to retention and storage configuration.

Cluster Observability Operator (COO)​

The Cluster Observability Operator (COO) is an optional OpenShift operator that enables administrators to create standalone monitoring stacks that are independently configurable. COO is ideal for users who need high customizability, scalability, and long-term data retention, especially in complex, multi-tenant enterprise environments.

COO can be used to:

  • Set up a highly available Prometheus instance that persists metrics
  • Enable remote writing of metrics to an Azure Monitor Prometheus workspace
  • Create independent monitoring stacks for different services and users

Azure Monitor integration​

Log forwarding: Starting from OpenShift Logging version 5.9, OpenShift supports native forwarding to Azure Monitor and Azure Log Analytics, available on clusters running OpenShift 4.13 or higher. This allows you to view and query the logs the platform and your workloads generate in Azure Monitor.

Metric remote write: COO enables remote writing of metrics to Azure Monitor Prometheus workspaces, allowing centralization of metrics in Azure.

The three pillars of observability​

Metrics: CPU, memory, network, and custom application metrics Logs: Application logs, container logs, and audit logs Traces: Distributed tracing for microservices

What to monitor​

RoleWhat to Monitor
DevelopersApplication errors, latency, request rate, pod restarts
Platform engineersNode utilization, cluster capacity, scheduling failures, API health
ArchitectsAvailability, dependency health, capacity, failure domains, cost

CI/CD and DevOps​

Production deployment pipeline​

Developer
↓
Git (source control)
↓
CI Pipeline (GitHub Actions / Azure DevOps)
↓
Build container image
↓
Security scan
↓
Push to Azure Container Registry
↓
CD Pipeline (GitHub Actions / Azure DevOps / OpenShift Pipelines)
↓
Deploy to OpenShift (oc apply / Helm / GitOps)
↓
Monitoring
↓
Rollback if required

OpenShift Pipelines​

OpenShift Pipelines is a Kubernetes-native CI/CD framework based on Tekton. It provides:

  • Declarative pipeline definitions as Kubernetes resources
  • Reusable tasks and steps
  • Integration with OpenShift's developer experience
  • CI/CD workflows that run inside the cluster

OpenShift GitOps​

OpenShift GitOps is based on ArgoCD and provides GitOps-style continuous delivery. It offers:

  • Declarative application delivery
  • Git as the single source of truth
  • Automatic sync between Git and cluster state
  • Least privileged access and version control

CI/CD integration options​

ToolIntegrationBest For
GitHub ActionsDirect, via OIDC or service principalGitHub-based workflows
Azure DevOpsNative via service connectionAzure-centric teams
OpenShift PipelinesNative, Kubernetes-basedTeams using OpenShift ecosystem
OpenShift GitOpsNative, ArgoCD-basedGitOps-driven teams

Infrastructure as Code​

  • Use Terraform for ARO cluster provisioning
  • Use Bicep or ARM templates for Azure infrastructure
  • Use Helm for application packaging and deployment
  • Store all IaC in version control

Operators​

The Operator pattern​

Operators are a Kubernetes-native way to automate the lifecycle of complex applications. The Operator pattern uses Custom Resources to define the desired state and controllers to reconcile the actual state with the desired state.

Custom Resource (desired state)
↓
Operator (controller)
↓
Kubernetes Resources (actual state)
↓
Running Application

Why Operators matter in OpenShift​

Operators are central to OpenShift's platform model. They provide:

  • Lifecycle automation – Install, upgrade, and manage applications
  • Day 2 operations – Backup, recovery, scaling, and configuration
  • Vendor integration – Database operators, middleware operators, platform components

Common Operators​

  • Database operators – PostgreSQL, MySQL, MongoDB, etc.
  • Middleware operators – Kafka, RabbitMQ, etc.
  • Platform operators – Ingress, monitoring, logging
  • Application operators – Custom operators for specific applications

When Operators add value​

Operators improve operational consistency when:

  • Deploying complex distributed systems
  • Managing stateful applications
  • Automating day-2 operations
  • Standardizing application lifecycle across environments

When Operators add complexity​

Operators can introduce additional complexity when:

  • Simple applications don't need them
  • The Operator itself requires significant management
  • Custom Operators are poorly maintained

OpenShift Developer Experience​

OpenShift Web Console​

The OpenShift web console provides a comprehensive interface for both administrators and developers:

  • Developer perspective – Create and deploy applications, view logs, manage routes
  • Administrator perspective – Manage cluster resources, projects, RBAC, and operators

OpenShift CLI (oc)​

oc extends kubectl with OpenShift-specific commands:

  • oc new-app – Create applications from source, images, or templates
  • oc new-project – Create projects with resource quotas
  • oc expose – Create routes to expose services
  • oc login – Authenticate to the cluster
  • oc whoami – Show current user

Developer workflow​

  1. Login – oc login with credentials
  2. Create or select project – oc new-project or oc project
  3. Deploy application – oc new-app or oc apply -f with YAML
  4. Expose application – oc expose service to create a route
  5. Monitor – oc logs, oc status, or use the web console
  6. Update – oc apply -f with updated YAML or oc patch

Focus areas​

Developers focus on:

  • Application code
  • Container images
  • Configuration (ConfigMaps, Secrets)
  • Deployment manifests

Platform teams provide:

  • Cluster infrastructure
  • Security and RBAC
  • Networking and ingress
  • Shared services

ARO for Microservices​

ARO is well-suited for enterprise microservices architectures.

Microservices architecture example​

External Client
↓
OpenShift Route (public ingress)
↓
API Gateway Service
↓
┌───────────────┬────────────────┬────────────────┐
↓ ↓ ↓ ↓
Orders Payments Users Notifications
Service Service Service Service
↓ ↓ ↓ ↓
└───────────────┴────────────────┴────────────────┘
│
▼
Service Bus (async events)
│
▼
Background Workers

Service boundaries​

Each microservice is deployed as a separate OpenShift Deployment:

  • Independent scaling based on workload
  • Independent deployment with rolling updates
  • Independent failure isolation
  • Own data store (database per service)

Communication patterns​

PatternImplementationUse Case
Synchronous HTTPOpenShift Services (ClusterIP)Request-response APIs
Async messagingAzure Service Bus / KafkaEvent-driven decoupling
Service meshOpenShift Service Mesh (Istio)Advanced traffic management

Independent deployment and scaling​

Each service has its own:

  • Deployment with rolling updates
  • Horizontal Pod Autoscaler
  • Service and Route configuration

Failure isolation​

  • Failures are contained within each service
  • Circuit breakers prevent cascading failures
  • Retries and timeouts protect downstream services

ARO for AI and Cloud-Native AI​

ARO is increasingly used for AI workloads, especially with native GPU support.

GPU support in ARO​

Azure Red Hat OpenShift now supports NVIDIA H100 and H200 GPU-based Azure Virtual Machine SKUs, enabling customers to run large-scale AI, machine learning, and high-performance computing (HPC) workloads. The NVIDIA H200 GPU SKU (Standard_ND96isr_H200_v5) provides 96 vCPUs.

OpenShift AI on ARO addresses GPU workload challenges with flexible, hourly GPU worker node availability, automated scale-out and scale-down capabilities, and a hybrid cloud strategy.

AI use cases on ARO​

  • AI inference APIs – Deploy trained models as scalable services
  • RAG applications – Host Retrieval-Augmented Generation APIs
  • AI agents – Deploy agentic AI services
  • MLOps platforms – Run ML pipelines and model training
  • Data processing – Process large datasets for AI/ML
  • Batch inference – Run inference on large batches of data

AI architecture example​

Client
↓
OpenShift Route
↓
AI API Service
↓
┌──────────────┬────────────────┬────────────────┐
↓ ↓ ↓ ↓
Azure OpenAI Azure AI Vector Store Data Store
Search

ARO vs. other platforms for AI​

PlatformBest For
AROEnterprise AI platforms, existing OpenShift skills, complex AI workloads, GPU workloads
Azure Container AppsSimpler AI APIs, event-driven AI, serverless scaling
Azure AI servicesManaged AI services (OpenAI, Vision, etc.)
AKSAzure-native Kubernetes, custom Kubernetes operators

When ARO is the right choice​

ARO is ideal when:

  • Your organization standardizes on OpenShift
  • You need enterprise Kubernetes for AI workloads
  • You require GPU acceleration with flexible availability
  • You have complex AI workloads requiring platform customization
  • You're building an enterprise AI platform

ARO and Enterprise Hybrid Cloud​

OpenShift consistency​

Organizations that use OpenShift on-premises and in other clouds can benefit from ARO's consistency:

  • Same OpenShift APIs and platform capabilities
  • Same developer experience across environments
  • Same operational patterns
  • Common tooling and skills

Hybrid cloud architecture​

On-Premises OpenShift
↕ (OpenShift APIs)
Azure Red Hat OpenShift
↕ (OpenShift APIs)
OpenShift on Other Clouds

Application portability​

OpenShift provides application portability across:

  • On-premises OpenShift
  • Azure Red Hat OpenShift
  • Red Hat OpenShift Service on AWS (ROSA)
  • Other OpenShift distributions

The architectural trade-off​

ApproachBenefitsTrade-offs
Standardize on OpenShiftConsistency, portability, common skillsOpenShift complexity, licensing costs
Standardize on Azure-native servicesAzure integration, lower cost, simplerVendor lock-in, different skill sets

ARO vs. Azure Kubernetes Service (AKS)​

DimensionAROAKS
PlatformOpenShift (enterprise Kubernetes)Vanilla Kubernetes
Vendor ecosystemRed Hat + MicrosoftMicrosoft + Kubernetes ecosystem
Kubernetes controlManagedManaged
OpenShift capabilities✅ Yes (operators, routes, SCCs, console)❌ No
Developer experienceOpenShift web console + oc CLIKubernetes dashboard + kubectl
OperatorsOpenShift ecosystem (Red Hat, certified, community)Kubernetes ecosystem
Azure integrationStrongStrong
Existing OpenShift skillsExcellent fitNot applicable
Azure-native KubernetesGoodExcellent
Security defaultsStricter (SCCs applied by default)Standard Kubernetes
SupportJoint Microsoft/Red HatMicrosoft
CostHigher (includes Red Hat licensing)Lower (free control plane)
Enterprise OpenShift standardizationExcellentNot applicable

Decision framework​

Choose ARO when:

  • The organization standardizes on OpenShift
  • Existing Red Hat/OpenShift expertise is important
  • Hybrid OpenShift consistency matters (on-prem + cloud)
  • OpenShift platform capabilities (routes, SCCs, operators) are required
  • You need joint Microsoft/Red Hat support

Choose AKS when:

  • Azure-native Kubernetes is preferred
  • The organization does not need OpenShift
  • Kubernetes ecosystem compatibility is sufficient
  • Azure integration and Kubernetes flexibility are primary goals
  • Cost is a primary concern

ARO vs. Self-Managed OpenShift​

AspectSelf-Managed OpenShiftARO
ControlFullLimited (platform managed)
Operational responsibilityHighLow
Infrastructure managementCustomer managesAzure manages
UpgradesCustomer plans and executesManaged by SRE team
AvailabilityCustomer responsibility99.95% SLA (with availability zones)
NetworkingFull controlVNet-integrated
SecurityCustomer responsibilityPlatform + customer shared
Platform maintenanceCustomer managesAutomated by platform
CostInfrastructure + Red Hat subscriptionIncluded in Azure billing
FlexibilityMaximumLimited to ARO capabilities

Why choose managed​

A managed service is preferable for organizations that want OpenShift without owning the full cluster lifecycle. The joint SRE team handles installation, scaling, security, monitoring, and updating of both control plane and worker nodes.

ARO vs. Azure Container Apps​

AspectAROAzure Container Apps
Abstraction levelKubernetes/OpenShift platformServerless container platform
Kubernetes/OpenShift access✅ Full OpenShift API❌ No Kubernetes access
Developer experienceOpenShift platformApplication-focused
Operational complexityHigherLower
ScalingHPA + Cluster AutoscalerKEDA-based, scale-to-zero
MicroservicesFull OpenShift supportBuilt-in service discovery
Platform controlHighLimited
NetworkingFull VNet controlManaged networking
Enterprise platformYesLimited

Decision framework​

Choose ARO when:

  • You need OpenShift
  • You need Kubernetes/OpenShift APIs and platform capabilities
  • Enterprise OpenShift standardization matters
  • You have advanced platform requirements
  • You need GPU support for AI workloads

Choose Container Apps when:

  • You want the simplest managed container application platform
  • You do not need Kubernetes/OpenShift control
  • You are building APIs, workers, or lightweight microservices
  • You want scale-to-zero

Cost Considerations​

Cost drivers​

  • Cluster infrastructure – VM costs for worker nodes
  • Red Hat licensing – Included in ARO pricing
  • Storage – Azure managed disks, Azure Files
  • Networking – Data transfer, load balancers
  • Monitoring – Log Analytics, Azure Monitor
  • Supporting Azure services – ACR, databases, etc.

ARO vs. AKS cost comparison​

ARO typically costs approximately 1.8x more than AKS due to the included Red Hat licensing. This premium provides:

  • OpenShift platform capabilities
  • Joint Microsoft/Red Hat support
  • Enterprise-grade security and compliance
  • Integrated CI/CD and monitoring

Cost optimization​

  • Right-size worker nodes based on workload requirements
  • Use autoscaling to match capacity with demand
  • Use spot instances for non-production workloads (where supported)
  • Separate development and production clusters to right-size each
  • Monitor and clean up unused resources

Production Best Practices​

Cluster design​

  • Use private clusters for production workloads
  • Deploy across availability zones for high availability
  • Use separate node pools for system and user workloads
  • Add infrastructure nodes for platform workloads
  • Right-size worker nodes based on workload characteristics

Identity and security​

  • Integrate Microsoft Entra ID for authentication
  • Use Azure Managed Identities instead of service principals
  • Apply least-privilege OpenShift RBAC
  • Use Security Context Constraints to enforce pod security
  • Enable audit logging
  • Use Azure Key Vault for secrets

Networking​

  • Plan VNet and subnet sizing carefully
  • Use private clusters for sensitive workloads
  • Configure egress appropriately (NAT Gateway or Azure Firewall)
  • Enable service endpoints for ACR access in private clusters
  • Use custom domains for production routes

Application management​

  • Use immutable image versions (avoid latest)
  • Scan container images for vulnerabilities
  • Use GitOps for declarative application delivery
  • Implement health probes (liveness, readiness, startup)
  • Use Pod Disruption Budgets for critical workloads

Observability​

  • Set up log forwarding to Azure Monitor
  • Configure metric remote write to Azure Monitor
  • Set up alerts for critical conditions
  • Monitor cluster and application health

Operations​

  • Test upgrades in staging before production
  • Back up cluster configuration and persistent data
  • Document dependencies and failure modes
  • Use Infrastructure as Code for cluster provisioning
  • Regularly review permissions and access

Common Mistakes​

Treating ARO as ordinary Kubernetes​

OpenShift adds significant capabilities and constraints beyond vanilla Kubernetes. Understand OpenShift-specific concepts like Projects, Routes, and Security Context Constraints.

Assuming ARO is identical to AKS​

ARO and AKS serve different purposes. ARO provides OpenShift with enterprise features; AKS provides vanilla Kubernetes. Choose based on your requirements.

Creating excessive clusters instead of using projects​

Projects provide logical isolation within a cluster. Creating separate clusters for every application increases cost and management complexity.

Using cluster-admin permissions for normal development​

Cluster-admin provides full access to the entire cluster. Use project-scoped permissions for developers.

Ignoring OpenShift RBAC​

Azure RBAC controls Azure resource management. OpenShift RBAC controls access to cluster resources. Both are needed.

Running privileged containers unnecessarily​

OpenShift applies restrictive SCCs by default. Don't grant privileged access unless absolutely required.

Storing secrets inside container images​

Secrets in images are exposed to anyone with image access. Use OpenShift Secrets or Azure Key Vault.

Using mutable image tags (latest) in production​

latest provides no versioning guarantees. Use immutable tags for production.

Ignoring resource requests and limits​

Without resource limits, workloads can consume all cluster resources. Always set requests and limits.

Ignoring cluster capacity​

Cluster creation requires at least 44 vCPUs. Monitor capacity and plan for growth.

Assuming pod autoscaling solves all capacity problems​

HPA scales pods but doesn't add nodes. Use Cluster Autoscaler for node scaling.

Exposing internal services publicly​

Use internal routes for services that shouldn't be publicly accessible.

Ignoring network egress​

Private clusters need configured egress for internet access. Plan egress routing.

Treating high availability as disaster recovery​

HA protects against node/zone failures. DR protects against region failures. Different strategies needed.

Keeping state inside pods​

Pods are ephemeral. Use PersistentVolumes or Azure services for persistent state.

Failing to monitor node and pod health​

Monitor both cluster and application health. Set up alerts for critical conditions.

Choosing ARO when a simpler managed container service is sufficient​

ARO adds complexity and cost. Consider Container Apps for simpler workloads.

Choosing ARO without considering existing organizational OpenShift skills​

ARO requires OpenShift knowledge. If the team doesn't have OpenShift experience, AKS or Container Apps may be better.

Underestimating OpenShift operational complexity​

ARO reduces platform management but doesn't eliminate it. You still need to manage workloads, networking, security, and operations.

Troubleshooting​

Cluster access failure​

  • Symptoms: Cannot access cluster API or console
  • Likely causes: Network connectivity, expired credentials, RBAC issues
  • Diagnose: Check network connectivity; verify authentication; check Azure RBAC
  • Fix: Use az aro list-credentials to retrieve credentials; verify network access

Authentication problems​

  • Symptoms: Cannot log in; unauthorized errors
  • Likely causes: Incorrect identity provider configuration; expired client secret
  • Diagnose: Check OpenShift OAuth configuration; verify Entra ID app registration
  • Fix: Update OAuth configuration; regenerate client secret

RBAC denial​

  • Symptoms: "Forbidden" errors when accessing resources
  • Likely causes: Missing RBAC permissions
  • Diagnose: Check OpenShift RBAC roles and bindings
  • Fix: Grant appropriate roles (view, edit, admin) at the project or cluster scope

Pod stuck in Pending​

  • Symptoms: Pod remains in Pending state
  • Likely causes: Insufficient cluster capacity; resource requests exceed available capacity
  • Diagnose: Check pod events (oc describe pod); check node capacity
  • Fix: Add nodes; reduce resource requests; use different node pool

Pod CrashLoopBackOff​

  • Symptoms: Pod crashes and restarts repeatedly
  • Likely causes: Application errors; missing dependencies; configuration issues
  • Diagnose: Check pod logs (oc logs); check pod events
  • Fix: Fix application code; correct configuration; increase resource limits

Image pull failure​

  • Symptoms: Pod shows ImagePullBackOff or ErrImagePull
  • Likely causes: Image doesn't exist; incorrect tag; ACR authentication failure
  • Diagnose: Check image name and tag; verify ACR access
  • Fix: Push the image; correct the tag; ensure pull secret is configured

Route unavailable​

  • Symptoms: Route doesn't respond to requests
  • Likely causes: Route configuration issues; service not running; TLS misconfiguration
  • Diagnose: Check route status (oc get routes); verify service endpoints
  • Fix: Correct route configuration; ensure service is running; update TLS certificates

Service-to-service communication failure​

  • Symptoms: Services can't communicate with each other
  • Likely causes: Incorrect service name; network policy blocking traffic
  • Diagnose: Verify service DNS resolution; check network policies
  • Fix: Use correct service names; configure network policies to allow traffic

Storage mount failure​

  • Symptoms: PVCs not binding; pod can't mount volumes
  • Likely causes: Storage class not available; insufficient permissions; quota exceeded
  • Diagnose: Check PVC status (oc get pvc); verify storage class
  • Fix: Use correct storage class; ensure sufficient quota; check permissions

High CPU or memory usage​

  • Symptoms: Nodes or pods using excessive resources
  • Likely causes: Application issues; insufficient resource limits; node under-provisioned
  • Diagnose: Check node and pod metrics; review application behavior
  • Fix: Increase resource limits; scale out; fix application performance issues

Autoscaling not behaving as expected​

  • Symptoms: HPA not scaling; scaling too aggressively or not at all
  • Likely causes: Metrics not available; incorrect configuration; insufficient cluster capacity
  • Diagnose: Check HPA status (oc describe hpa); verify metrics-server
  • Fix: Correct HPA configuration; ensure metrics are available; add cluster capacity

Practical Learning Path​

  1. Learn containers and Docker – Basic container concepts, Dockerfiles, image building

  2. Learn Kubernetes fundamentals – Pods, Deployments, Services, Namespaces, ConfigMaps, Secrets

  3. Understand OpenShift concepts – Projects, Routes, Security Context Constraints, Operators

  4. Create an ARO cluster – Use Azure CLI or portal; understand prerequisites (vCPUs, subnets)

  5. Explore the OpenShift Console – Developer and Administrator perspectives

  6. Deploy a containerized application – Using oc new-app or YAML manifests

  7. Learn Projects and RBAC – Create projects; assign roles (view, edit, admin)

  8. Configure Routes and Services – Expose applications externally with routes

  9. Configure persistent storage – PVs and PVCs for stateful applications

  10. Integrate Azure Container Registry – Push images; configure pull secrets

  11. Integrate Microsoft Entra ID – Configure OIDC identity provider

  12. Configure autoscaling – HPA for pods; Cluster Autoscaler for nodes

  13. Implement CI/CD – OpenShift Pipelines or GitOps

  14. Add monitoring and logging – COO for custom monitoring; log forwarding to Azure Monitor

  15. Secure workloads – SCCs, network policies, secrets management

  16. Build a microservices architecture – Multiple services, internal communication, routes

  17. Explore Operators – Understand Operator pattern; deploy common operators

  18. Design hybrid OpenShift architectures – Connect ARO with on-premises OpenShift

  19. Evaluate ARO vs. AKS – Understand the decision framework

  20. Apply ARO to AI/cloud-native workloads – GPU nodes, model inference, RAG

Interview and Architecture Questions​

What is Azure Red Hat OpenShift? Azure Red Hat OpenShift is a fully managed, enterprise-grade Kubernetes platform jointly developed, operated, and supported by Red Hat and Microsoft on the Azure cloud. It provides OpenShift capabilities without the complexity of managing the underlying infrastructure.

How is ARO different from AKS? ARO provides OpenShift with enterprise features (routes, SCCs, operators, web console) and joint Microsoft/Red Hat support. AKS provides vanilla Kubernetes with Azure-native integrations and lower cost. ARO typically costs more due to Red Hat licensing.

What does Microsoft manage in ARO? Microsoft and Red Hat manage the installation, scaling, security, monitoring, and updating of both control plane and worker nodes. The joint SRE team ensures high availability and resilience.

What does the customer manage? Customers manage worker node configuration, scaling, application workloads, networking (VNet), identity integration, storage configuration, and monitoring setup.

What is the difference between Kubernetes and OpenShift? OpenShift builds on Kubernetes with an enterprise-ready platform including container management, automation, networking, CI/CD, monitoring, registry, and authentication—all tested together. Kubernetes is CaaS; OpenShift/ARO is PaaS.

What is an OpenShift Project? A Project is OpenShift's abstraction over Kubernetes namespaces with additional features: resource quotas, limit ranges, project-scoped RBAC, and network policies.

What is an OpenShift Route? A Route is OpenShift's native ingress solution that exposes services externally with hostname-based routing, path-based routing, TLS termination, and traffic splitting.

What is an Operator? An Operator is a Kubernetes-native pattern for automating the lifecycle of complex applications using Custom Resources and controllers. Operators manage installation, upgrades, backup, and recovery.

How does ARO integrate with Azure networking? ARO is deployed into a customer's Azure VNet with dedicated subnets for master and worker nodes. It supports private clusters, service endpoints, and Private Link.

How does identity work in ARO? ARO integrates with Microsoft Entra ID as an OIDC identity provider. Users authenticate via Entra ID; group claims map Entra ID groups to OpenShift RBAC. Managed Identities and Workload Identity are supported for platform operators.

How do Azure RBAC and OpenShift RBAC differ? Azure RBAC controls who can manage ARO resources in Azure. OpenShift RBAC controls who can access cluster resources (projects, pods, services). Both are needed for comprehensive access control.

How would you secure an ARO cluster? Use private clusters, integrate Entra ID, apply least-privilege RBAC, use SCCs, enable audit logging, use network policies, scan images, use Azure Key Vault for secrets, and monitor regularly.

How would you design a production ARO architecture? Use private clusters across availability zones, separate node pools for system and user workloads, add infrastructure nodes, integrate Entra ID, use ACR with private endpoint, implement monitoring with Azure Monitor, and use GitOps for deployment.

When would you choose ARO instead of AKS? Choose ARO when your organization standardizes on OpenShift, needs OpenShift capabilities, wants hybrid OpenShift consistency, or requires joint Microsoft/Red Hat support.

When would you choose Container Apps instead of ARO? Choose Container Apps when you want the simplest managed container platform, don't need Kubernetes/OpenShift control, and are building APIs or lightweight microservices.

How would you design hybrid OpenShift? Use OpenShift APIs for consistency across on-premises OpenShift and ARO. Use the same application definitions, RBAC patterns, and deployment processes. Use Azure networking (VPN, ExpressRoute) for connectivity.

How would you deploy AI workloads on ARO? Use GPU-enabled worker nodes (NVIDIA H100/H200). Deploy model-serving containers. Use OpenShift AI or custom Operators for ML pipelines. Integrate with Azure AI services and vector stores.

How would you troubleshoot a pod stuck in Pending? Check pod events (oc describe pod). Check node capacity and resource availability. Verify resource requests don't exceed available capacity. Check for node taints or tolerations.

How would you troubleshoot an image pull failure? Check the image name and tag. Verify the image exists in ACR. Check pull secret configuration. For private clusters, verify ACR service endpoint is enabled.

Key Takeaways​

Azure Red Hat OpenShift (ARO) is a fully managed, enterprise-grade Kubernetes platform jointly developed, operated, and supported by Red Hat and Microsoft on Azure.

ARO provides OpenShift's enterprise capabilities on Azure – routes, SCCs, operators, built-in CI/CD, monitoring, and a developer-friendly web console.

The platform is managed by a joint SRE team from Red Hat and Microsoft, handling installation, scaling, security, monitoring, and updates. This significantly reduces operational burden.

ARO is built on Azure infrastructure – VMs, network security groups, and storage accounts are deployed into your Azure subscription. You integrate with Azure VNets, subnets, and Azure services.

Private clusters require careful networking configuration, including service endpoints for ACR access and custom egress routing via NAT Gateway or Azure Firewall.

Identity integrates with Microsoft Entra ID as an OIDC provider. Managed Identities and Workload Identity provide secret-free authentication for platform operators.

OpenShift adds Security Context Constraints (SCCs) that enforce stricter security defaults than vanilla Kubernetes. Pods require explicit opt-in for privileged operations.

ARO supports NVIDIA H100 and H200 GPUs for large-scale AI, machine learning, and HPC workloads. OpenShift AI provides flexible GPU worker node availability.

Observability integrates with Azure Monitor – log forwarding is native in OpenShift 4.13+, and the Cluster Observability Operator enables remote metric writes.

Choose ARO when your organization standardizes on OpenShift, needs OpenShift capabilities, or requires hybrid OpenShift consistency. Choose AKS for Azure-native Kubernetes without OpenShift.

Production best practices include: use private clusters, deploy across availability zones, integrate Entra ID, apply least-privilege RBAC, use immutable image versions, implement monitoring, and use GitOps for declarative deployment.