Lyft has successfully migrated hundreds of its production Apache Flink jobs from a proprietary Kubernetes operator to the Apache Flink Kubernetes Operator. This change was detailed in an August 31 Lyft Engineering post by Maheep Myneni, Arda Kuyumcu, and Prem Santosh Udaya Shankar. The move was prompted by challenges with the in-house operator, including difficulties with Flink upgrades and limitations in autoscaling and rollback functionalities.
The in-house operator, developed by Lyft in 2020, lacked a dedicated control plane for Flink on Kubernetes. Its dual-deployment upgrade process involved starting a new cluster, taking a savepoint, canceling the old job, and restoring from the savepoint. This method was prone to failures due to a lack of retry logic and idempotency in the savepoint trigger. Additionally, memory reservation for non-JVM components was managed by a single knob, leading to either out-of-memory errors or wasted resources.
The Apache Flink Kubernetes Operator offers several improvements. It treats last-state as a primary upgrade mode, allowing restoration from high-availability metadata or the latest checkpoint even if the JobManager is unhealthy. Lyft implemented a translation layer to convert their legacy FlinkApplication specifications to FlinkDeployment resources, ensuring compatibility. This layer mapped jarName to jarURI, moved node selectors into PodTemplateSpec objects, injected environment variables, and defaulted upgrades to last-state.
The Apache operator initiates JobManagers first to manage TaskManager lifecycles. The migration also replaced the legacy dual deployments with stop-then-start deployments, which initially caused 3 to 6 minutes of downtime for typical deployments and about 20 minutes for larger jobs. To mitigate this, Lyft adopted FlinkBlueGreenDeployment, a Custom Resource Definition (CRD) that allows new versions to run alongside old ones before cutover. Lyft also contributed an upstream fix for a configuration-rename bug (FLINK-38548) encountered during their testing of the BlueGreen deployment feature.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Lyft has transitioned hundreds of production Apache Flink jobs from its custom-built Kubernetes operator to the open-source Apache Flink Kubernetes Operator. This migration addresses issues with maintenance, autoscaling, and rollback capabilities present in their previous in-house solution, enabling features like last-state upgrades and in-place autoscaling.