← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

Lyft Migrates Hundreds of Flink Jobs to Apache Flink Kubernetes Operator

🔄 Updated 6d ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Lyft moved hundreds of Flink jobs to the Apache Flink Kubernetes Operator.
  • The in-house operator had issues with upgrades, autoscaling, and rollbacks.
  • The new operator enables last-state upgrades and in-place autoscaling.
  • Lyft contributed a fix for a BlueGreen deployment bug to the Apache operator.

Migration to Apache Flink Kubernetes Operator

Lyft has successfully migrated hundreds of its production Apache Flink jobs from a proprietary Kubernetes operator to the Apache Flink Kubernetes Operator. This change was detailed in an August 31 Lyft Engineering post by Maheep Myneni, Arda Kuyumcu, and Prem Santosh Udaya Shankar. The move was prompted by challenges with the in-house operator, including difficulties with Flink upgrades and limitations in autoscaling and rollback functionalities.

Addressing Legacy Operator Limitations

The in-house operator, developed by Lyft in 2020, lacked a dedicated control plane for Flink on Kubernetes. Its dual-deployment upgrade process involved starting a new cluster, taking a savepoint, canceling the old job, and restoring from the savepoint. This method was prone to failures due to a lack of retry logic and idempotency in the savepoint trigger. Additionally, memory reservation for non-JVM components was managed by a single knob, leading to either out-of-memory errors or wasted resources.

Benefits of the Apache Operator

The Apache Flink Kubernetes Operator offers several improvements. It treats last-state as a primary upgrade mode, allowing restoration from high-availability metadata or the latest checkpoint even if the JobManager is unhealthy. Lyft implemented a translation layer to convert their legacy FlinkApplication specifications to FlinkDeployment resources, ensuring compatibility. This layer mapped jarName to jarURI, moved node selectors into PodTemplateSpec objects, injected environment variables, and defaulted upgrades to last-state.

Improved Deployment and Contributions

The Apache operator initiates JobManagers first to manage TaskManager lifecycles. The migration also replaced the legacy dual deployments with stop-then-start deployments, which initially caused 3 to 6 minutes of downtime for typical deployments and about 20 minutes for larger jobs. To mitigate this, Lyft adopted FlinkBlueGreenDeployment, a Custom Resource Definition (CRD) that allows new versions to run alongside old ones before cutover. Lyft also contributed an upstream fix for a configuration-rename bug (FLINK-38548) encountered during their testing of the BlueGreen deployment feature.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~26 min · 21 stories · Sep 23

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Lyft has transitioned hundreds of production Apache Flink jobs from its custom-built Kubernetes operator to the open-source Apache Flink Kubernetes Operator. This migration addresses issues with maintenance, autoscaling, and rollback capabilities present in their previous in-house solution, enabling features like last-state upgrades and in-place autoscaling.