Lester Martin

Lester Martin

dev advocate, trainer, curriculum developer, blogger, data engineer

Atlanta, Georgia, United States

Actions

Lester Martin is a seasoned developer advocate, trainer, blogger, data engineer, and polyglot programmer focused on data pipelines & data lake analytics using Trino, Iceberg, Hive, Spark, Flink, Kafka, NiFi, NoSQL databases, and, of course, classical RDBMSs. Learn more about Lester at https://linktr.ee/lestermartin.

Area of Expertise

  • Information & Communications Technology

Topics

  • devrel
  • Software Development
  • Data Lake
  • Data Engineering
  • Data Lakehouse
  • Trino
  • Starburst
  • Apache Iceberg
  • Apache Hive
  • Spark
  • PySpark
  • Apache Kafka
  • apache nifi

Apache Iceberg Deep Dive Workshop

Get hands-on with Apache Iceberg in this introductory workshop. We’ll cover the full table lifecycle, from creating your first Iceberg table to managing snapshots, performing rollbacks, and running compactions to improve performance.

In this workshop, you’ll learn how to:

- Create and manipulate Iceberg tables with familiar SQL
- Use snapshots and time travel to audit changes and reproduce results
- Perform rollbacks to recover quickly from bad writes or deployments
- Run compaction and maintenance to improve performance and reduce small files
- Understand the table lifecycle, schema/partition evolution, metadata, and governance touchpoints

Apache Iceberg ingestion with Apache NiFi

A cornerstone requirement of an Icehouse (Iceberg + Trino) is data ingestion. One approach is to leverage Apache NiFi. NiFi, a multimodal data pipelining tool, has a multitude of processors that can be assembled into a flow to address your specific scenarios. NiFi's low-code/no-code approach allows data engineers to rapidly build, deploy, and monitor their data ingestion & transformation pipelines. NiFi also allows custom processor development with a variety of languages, including Java and Python.

This presentation will iterate through a few common approaches and ultimately demonstrate a rich data pipeline that sources data from Kafka, performs typical transformation processing (including enrichment), and loads data into a high-performance Iceberg table that will be consumed via Trino.

Understanding & Exploring Apache Iceberg v3

Apache Iceberg is an open-source table format that provides database-like functionalities such as ACID transactions, schema evolution, and time travel for large analytic datasets stored in data lakes. Iceberg Spec v3 introduces significant advancements, including binary deletion vectors for faster deletes and updates, richer data types like variant for semi-structured data, nanosecond-precision timestamps, and built-in row lineage for enhanced data governance.

In addition to learning what these cool features can do for you, you'll see a live demo of many of the popular features and walk away with a hands-on exercise in case you want to learn by doing, too.

Ibis: Bringing Optionality to Python Dataframes

Love the power of writing lazy executed Dataframe code in Python that runs on your favorite distributed data cluster? Would like some flexibility to swap out your processing engine for another? If so, you need Optionality in your Python Dataframe API.

Ibis, https://ibis-project.org/, offers a Python Dataframe API that lets your code run on nearly 20 backend data processing systems. It is THE portable Dataframe library. Imagine being able to run your Ibis code in Polars on your laptop and then moving it to PySpark in your favorite cloud provider with just changing a property. No need to imagine; you can do it today.

This presentation walks you through the features available in Ibis as well as compares it with other popular Dataframe APIs. You'll see how to mix-and-match SQL and Dataframe API transformations as desired and how to change the backend system were your code is executed.

You will see a demo of a job running in DuckDB for local testing and then with a single line of code being changed run in a Trino cluster. Step-by-step instructions will be provided to follow along on your laptop or to run the exercise yourself later.

Optimizing Your Apache Iceberg Lakehouse

Join Lester, author for the upcoming O'Reilly book Optimizing Your Apache Iceberg Lakehouse, for a practical, fast-paced session on improving query performance across your data lakehouse. While we focus on Apache Iceberg, the techniques apply broadly to Delta Lake and Apache Hive as well.

We’ll start with optimizations you can apply today as a table consumer: maintaining statistics, using effective filtering and projection, and leveraging caching to reduce latency.

Then we will go under the hood to show how your lakehouse tables should be structured and maintained to improve performance at scale, covering join optimization and file size considerations, as well as compaction, partitioning, bucketing, and file-level sorting.

You’ll learn how to:

- Reduce the amount of scanned data and speed up queries with statistics, filtering, and projection pruning.

- Design tables for scale with partition strategies based on best practices.

- Maintain tables with compaction, metadata rewriting, and expiration.

You will leave with practical guidance you can apply immediately—no replatforming required.

Early releases of the book available at https://learning.oreilly.com/library/view/optimizing-your-apache/0642572327040/

Michigan Technology Conference 2026 Sessionize Event Upcoming

October 2026 Rochester, Michigan, United States

DataEngBytes - Sydney

Building Trino data pipelines with SQL or Python
Implementing the medallion architecture with Starburst

July 2025 Sydney, Australia

Community Over Code Asia 2025 Sessionize Event

July 2025

DataEngBytes - Melbourne

Building Trino data pipelines with SQL or Python
Implementing the medallion architecture with Starburst

July 2025 Melbourne, Australia

Berlin Buzzwords 2025

Apache Iceberg ingestion with Apache NiFi
https://www.youtube.com/watch?v=2yH9PfiXb9Y

June 2025 Berlin, Germany

Lester Martin

dev advocate, trainer, curriculum developer, blogger, data engineer

Atlanta, Georgia, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top