~/devreads

#data engineering

23 posts

20 Jul

Deepika Saini 5 min read

“We have multiple dashboards showing different numbers. Which one is correct?” If you’ve worked in data long enough, you’ve probably heard this question more times than you’d like. Sales reports one revenue figure. Finance reports another. Product Analytics has a third. Executives spend more time debating whose dashboard is correct than discussing what action to take. The natural response is…

leadershipdatabricksdata-engineeringdata-architecturedata-strategy

13 Jul

Netflix Technology Blog 23 min read

By Parth Jain , Rakesh Sukumar , Yingwu Zhao , Renzo Sanchez-Silva & Nathan Fisher A deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale: from streaming architectures and distributed aggregation pipelines to time-travel queries and the methodology that made it work. Introduction In our first post , we introduced the problem: engineers…

backend-developmentdistributed-systemsdata-engineeringobservabilitysoftware-engineering

30 Jun

Sagibhuvana 7 min read

Expedia Group Technology — Innovation Using large language models to reveal bottlenecks in Spark SQL execution plans Photo by Luis del Río If you’ve ever stared at a 300-plus-node physical plan at 2 a.m. trying to spot a missing broadcast or one cursed skewed partition, this is for you. Spark makes it deceptively easy to write complex SQL that looks…

apache-sparkbig-datainnovationllmdata-engineering

18 Jun

Manav Mehta 7 min read

Building a Centralized Alerting Framework for Data Quality using Snowflake Before We Knew Better As our data platform grew, so did the number of pipelines, scheduled tasks, and data quality checks running every day. While Snowflake provided a reliable platform for storing and processing data, operational monitoring was fragmented across multiple systems. Data quality failures were often discovered only after…

snowflakeincident-managementdata-qualitydata-engineeringobservability

11 Jun

9 Jun

Patrick Lam 9 min read

How Airbnb’s data engineers and analytics engineers built a consistent and flexible data modeling framework to support the expansion into Homes, Experiences, and Services. By : Patrick Lam , Namrata Lamba , Jamie Stober With the May 2025 Summer Release, Airbnb redesigned its app, relaunched Experiences, and debuted Services, pushing us beyond our traditional Homes focus. For the data teams,…

data-engineeringanalytics-engineeringtechnologydata-modelingdata-architecture

3 Jun

Poorva Patil 6 min read

Photo by Corinne Kutz on Unsplash Before we knew better Our orchestration system started as a simple internal solution to manage event pipelines and trigger downstream jobs. Over time, as more workflows and dependencies were added, it gradually evolved into a tightly coupled monolithic scheduler that became increasingly difficult to understand and maintain. Understanding how a workflow executed often meant…

etlapache-airflowawsdata-engineeringsoftware-architecture

25 May

Aarav Nigam 7 min read

Authors: Charan , Aarav Nigam Special thanks to Meghana Negi for her contribution and guidance throughout the project. Introduction In hyperlocal delivery, finding a customer’s location is only half the problem. A latitude-longitude pin can tell us where a delivery ends on the map, but not how a delivery partner should interpret that location in the real world. In dense…

point-of-interestdata-engineeringhyperlocal-deliverygeospatial-dataaddress-resolution

5 May

Mahendran Vasagam 13 min read

Excerpt By 2024, Slack’s data platform had accumulated 700+ SSH-based operators orchestrating critical data pipelines. We’re talking daily search indexing that processed terabytes of data, analytics jobs powering business intelligence, the whole shebang. Every single one of these jobs required direct SSH access to production AWS Elastic MapReduce (EMR) clusters. We had a massive security…

uncategorizedairflowawsbig-datadata-engineering

22 Jan

Abhishek Bharti 6 min read

Flipkart serves thousands of sponsored ads across various pages and with search queries. Advertisers are charged based on the impression views or clicks generated by their campaign content. In the high-velocity domain of AdTech, the latency between an ad impression/click and a budget deduction represents a direct financial risk . If the system lags, advertisers overspend; if it blocks, revenue…

apache-flinkadtechlambda-architecturestream-processingdata-engineering

2 May 2025

Sameeksha Bhatia 7 min read

Load Testing API’s on Redshift & Snowflake — A Quick POC Overview At Helpshift, our data platform follows a Lakehouse architecture , combining the best of both data lakes and data warehouses . This architecture allows us to store and analyze large amounts of raw data in a structured and organized manner, while also providing the scalability and low-cost storage…

load-testingdata-engineeringsnowflakeredshiftperformance

2 Jul 2024

Nilanjana Mukherjee 9 min read

Slack Data Engineering recently underwent data workload migration from AWS EMR 5 (Spark 2/Hive 2 processing engine) to EMR 6 (Spark 3 processing engine). In this blog, we will share our migration journey, challenges, and the performance gains we observed in the process. This blog aims to assist Data Engineers, Data Infrastructure Engineers, and Product…

uncategorizedanalyticsawsbig-datadata-engineering

8 May 2024

Lakshmi Mohan 8 min read

The Data Engineering team is responsible for Slack’s data lake, analytics dashboards, and other data services. The team’s mission is to empower users to leverage data to make decisions quickly, accurately, and easily. Slack’s data lake grew in size from sub-petabyte to over 100 petabytes in recent years and it now spans millions of tables.…

data-engineering

8 Jan 2024

Bisman Sodhi 4 min read

Hi my name is Bisman and I studied Computer Science at University of California, Santa Barbara. During summer of 2022, I had the most amazing experience working as a Software Engineer Intern on Strava’s Data Platform Team. In the first fews weeks, I learned the tools my team uses and then spent the rest of the time working on my…

software-engineeringdata-platformsdata-engineering

9 Oct 2023

28 Apr 2023

Lou Kratz 7 min read

(cover image from ThisisEngineering RAEng) Let’s face it: software is easier to write than maintain. This is why we, as software engineers, prefer to just “rip it out and start over” instead of trying to understand what another developer (or our past self) was thinking. We seem to have collectively forgotten that “programs must be […]

uncategorizedartificial intelligenceawsaws sagemakerdata engineering

13 Sept 2022

17 Aug 2021

Samuel Bock 8 min read

Reinventing how the world does work inevitably creates a lot of data. Each year, Slack’s scale has increased and the volume of data ingested and stored has kept pace. To make it possible to understand relationships within our data, we’ve invested heavily in an automated data lineage framework. This facilitates producer/consumer coordination, improves risk mitigation,…

uncategorizedbig-datadata-engineering

28 Jul 2021

Sarah Henkens 10 min read

With the release of Slack Connect, people can now collaborate both with internal employees and external organizations in the same channel. To make this as smooth as possible, Slack does predictive email analysis to classify and recommend the best way for a user to work with people they want to collaborate with. To accomplish this,…

uncategorizedalgorithmsdata-engineeringinfrastructure

29 Apr 2021

16 Nov 2020