Ido Nadler
Turn your architecture on its head, and achieve huge improvements!
Actions
Ido is a big data team lead at Nielsen. His work focuses on building massive data pipelines (~250 Billion events/day) and infrastructure for running machine learning algorithms. Ido's projects run on AWS using a variety of technologies like Kafka, Spark, Airflow, Kubernetes, and more. He likes to continuously experiment with new technologies, tackle challenging problems, and find those better, more elegant, and cost-effective solutions.
Should you read Kafka as a stream or in batch? Should you even care?
Should you consume Kafka in a stream OR batch? When should you choose each one? What is more efficient, and cost effective?
In this talk we’ll give you the tools and metrics to decide which solution you should apply when, and show you a real life example with cost & time comparisons.
To highlight the differences, we’ll dive into a project we’ve done, transitioning from reading Kafka in a stream to reading it in batch.
By turning conventional thinking on its head and reading our multi-petabyte Kafka stream in batch using Spark and Airflow, we’ve achieved a huge cost reduction of 65% while at the same time getting a more scalable and resilient solution.
We’ll explore the tradeoffs and give you the metrics and intuition you’ll need to make such decisions yourself.
We’ll cover:
Costs of processing in stream compared to batch
Scaling up for bursts and reprocessing
Making the tradeoff between wait times and costs
Recovering from outages
And much more…
Scaling your Kafka streaming pipeline can be a pain - but it doesn’t have to be!!
Kafka data pipeline maintenance can be painful.
It usually comes with complicated and lengthy recovery processes, scaling difficulties, traffic ‘moodiness’, and latency issues after downtimes and outages.
It doesn’t have to be that way!
We’ll examine one of our multi-petabyte scale Kafka pipelines, and go over some of the pitfalls we’ve encountered. We’ll offer solutions that alleviate those problems, and go over comparisons between the before and after . We’ll then explain why some common sense solutions do not work well and offer an improved, scalable and resilient way of processing your stream.
We’ll cover:
• Costs of processing in stream compared to in batch
• Scaling out for bursts and reprocessing
• Making the tradeoff between wait times and costs
• Recovering from outages
• And much more…
Kafka Summit London 2022 Sessionize Event
Kafka Summit Americas 2021 Sessionize Event
Kafka Summit APAC 2021 Sessionize Event
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top