TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Articles/KEDA Turns Kubernetes Autoscaling Into a Queue-Depth Problem
Articles

KEDA Turns Kubernetes Autoscaling Into a Queue-Depth Problem

A CNCF Blog guide shows why event-driven Kubernetes workers need SQS queue depth as the autoscaling signal, not CPU or memory utilization, and how KEDA and EKS Pod Identity make it work.

August 1, 2026 5 Min Read
54

A pod that spends the night at 4 percent CPU can still be failing its job. A CNCF Blog post published July 31, 2026 by Albena Galabova of Itgix walks through why that mismatch is common in event-driven Kubernetes workloads, and how KEDA, the Cloud Native Computing Foundation’s graduated event-driven autoscaler, replaces CPU and memory with the one number that actually tracks the work waiting to be done: how many messages are sitting in an Amazon SQS queue.

Table Of Content

  • Why CPU and Memory Autoscaling Miss the Point
  • KEDA Extends the HPA With a Backlog Metric
  • Reading the ScaledObject
  • Identity Without an OIDC Provider
  • What Actually Breaks in Production
  • A Pattern Bigger Than SQS

The pattern shows up anywhere a Kubernetes worker pulls jobs off a queue instead of answering requests directly: order processing, image or video transcoding, webhook and notification delivery, batch inference. All of those workloads can look calm on a CPU or memory dashboard while a queue quietly backs up behind them, because the pod spends most of its time waiting on network I/O, not computing.

Why CPU and Memory Autoscaling Miss the Point

The Kubernetes Horizontal Pod Autoscaler (HPA) was built around resource utilization: watch a Deployment’s CPU or memory, add replicas when it climbs, remove them when it falls. That model fits a request-driven web service well. It fits a queue worker badly. A worker pulling messages from SQS can sit almost idle on CPU while thousands of messages back up behind it, since most of its time goes to waiting on the network, not computing. The reverse failure mode shows up too: a pod can still look busy on CPU metrics well after the real backlog has already cleared. As the CNCF post frames it, in event-driven Kubernetes architectures backlog, not infrastructure utilization, is the signal that actually reflects system pressure.

KEDA Extends the HPA With a Backlog Metric

KEDA (Kubernetes-based Event Driven Autoscaling) does not replace the HPA. Its own GitHub README describes it as a Kubernetes Metrics Server that integrates natively with components such as the HPA and carries no external dependencies, and states plainly that KEDA is a Cloud Native Computing Foundation graduated project. In practice that means KEDA supplies the HPA with a metric it would not otherwise have access to, an SQS queue’s backlog, and then lets the HPA run the same scaling math it was already built to do.

The request flow the CNCF post lays out is short: producers write messages to an SQS queue, KEDA polls the queue’s attributes to read the backlog, KEDA updates the metric the HPA reads, the HPA scales the worker Deployment, Kubernetes schedules the additional pods, and replicas scale back down as the queue drains. Every decision in that loop traces back to one number: how many messages are outstanding right now.

Reading the ScaledObject

KEDA’s aws-sqs-queue trigger is configured through a ScaledObject custom resource. The example architecture in the CNCF post sets pollingInterval: 10, cooldownPeriod: 120, minReplicaCount: 0, and maxReplicaCount: 30, then targets queueLength: "10" with activationQueueLength: "1": each pod is sized to handle 10 messages, and the deployment activates from zero as soon as a single message arrives. KEDA’s own scaler documentation confirms queueLength defaults to 5 and activationQueueLength defaults to 0 when a ScaledObject leaves them unset, so both values above are deliberate tuning choices rather than defaults left in place.

The replica math itself is not just ApproximateNumberOfMessages. By default KEDA adds ApproximateNumberOfMessagesNotVisible, the SQS attribute for messages currently in flight, since a message a consumer has not yet deleted still represents outstanding work. That behavior is controlled by a scaleOnInFlight flag (true by default); a separate scaleOnDelayed flag, false by default, can fold delayed messages into the same count too. Getting either flag wrong for a given workload produces two of the specific failure modes the CNCF post’s own troubleshooting section names: counting in-flight messages when a consumer’s visibility timeout runs long can make KEDA see phantom backlog and over-scale, while a queue that never clears its delayed or unacknowledged messages can keep a deployment from ever scaling back to zero.

Identity Without an OIDC Provider

The worker pod still needs real AWS permissions to call GetQueueAttributes, GetQueueUrl, ReceiveMessage, DeleteMessage, and ChangeMessageVisibility against the queue. The CNCF post authenticates that pod with EKS Pod Identity rather than the older IAM Roles for Service Accounts (IRSA) pattern it lists as the alternative: a TriggerAuthentication resource with podIdentity.provider: aws tells KEDA to pick up credentials the same way the worker’s own service account already does.

AWS’s documentation lays out why Pod Identity has become the simpler default: it maps an IAM role straight to a Kubernetes service account without standing up an OIDC identity provider for the cluster first, it reuses one IAM principal (pods.eks.amazonaws.com) instead of a separate trust relationship per cluster, and it lets a single EKS Pod Identity Agent on each node issue credentials once per node instead of once per pod. A cluster can register up to 5,000 of these associations. KEDA still exposes an older identityOwner parameter for choosing whether SQS permissions come from the pod’s own identity or from the KEDA operator’s identity, but its own scaler docs mark that field deprecated as of KEDA v2.13 and slated for removal in v3, which makes the explicit TriggerAuthentication and Pod Identity path used above the one with a future, not just the one with less setup today.

What Actually Breaks in Production

The CNCF post’s own verification step is simple: send 50 messages to the queue, watch kubectl get scaledobject and kubectl get pods -w, and confirm replicas actually return to zero once the queue empties, before trusting the setup with real traffic. Its troubleshooting table reads like a list of the ways that confidence can turn out to be premature. No scale-out at all usually traces back to a wrong queue URL or a missing IAM permission, and the fix is boring but effective: kubectl describe scaledobject plus the KEDA operator’s own logs. An HPA that exists but never scales anything is often an authentication failure or a metric-fetch error rather than a scaling-math problem.

The tuning advice that follows is just as unglamorous. Pick queueLength from measured throughput instead of a guess. Tune cooldownPeriod and the HPA’s own behavior.scaleDown settings separately, since they govern different scale-down paths; the CNCF example pairs a 120-second cooldown with a 60-second HPA stabilizationWindowSeconds. Turn on fallback replicas so a metrics outage degrades gracefully instead of freezing the deployment at whatever replica count it last saw. KEDA’s own scaling documentation adds a lower-level version of the same discipline: an optional metrics-caching feature can serve the HPA’s frequent metric requests (roughly every 15 seconds, Kubernetes’ own default sync period) from a cache refreshed only once per pollingInterval, so the scaler service itself is not hammered faster than the underlying number actually changes.

A Pattern Bigger Than SQS

The CNCF post is explicit that none of this is SQS-specific. KEDA ships scalers for Kafka, RabbitMQ, Azure Service Bus, Google Cloud Pub/Sub, and dozens of other event sources built on the same shape: read a backlog number from somewhere KEDA does not manage, hand it to the HPA as a metric, and let Kubernetes scale on demand instead of on a proxy for demand. The post’s closing point is the more durable lesson underneath the SQS example: the hard part was never installing KEDA, it was picking a backlog signal that actually represents the work still waiting, then tuning cooldowns, activation thresholds, and in-flight accounting to match how that specific queue and that specific worker behave. Get the signal right, and the HPA does the rest of the job it was always designed to do.

Tags:

AutoscalingAWSCloud NativeKEDAKubernetes

Share

Two AMD Epyc 9754 processors installed in a dual-socket server motherboard
Previous Post

Nvidia Details Vera CPU’s Custom Olympus Cores in a Direct Challenge to Intel and AMD

A hand-drawn UX wireframe sketch on grid paper labeling a page's Logo, Profile, CTA button, and Details sections
Next Post

How to Build a SaaS Landing Page With Next.js and shadcn/ui

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

Blue-lit server racks in a modern data center, illustrating the compute infrastructure behind the AI boom.
Articles

The AI Boom Is Spending Real Money Before Proving Real Returns

June 7, 2026
Technician working with a laptop beside server racks, representing enterprise AI retrieval infrastructure
Articles

Google’s Agentic RAG Push Makes Enterprise AI Less of a One-Shot Guess

June 7, 2026
A person with a laptop and smartphone, representing digital attention and AI-assisted work
Articles

AI Chatbots Are Making Attention a Design Problem

June 7, 2026
A customer-support representative wearing a headset against a dark studio background.
Articles

The Meta AI Support Hack Was a Plain Old Authorization Failure

June 7, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026