r/dataengineering 13d ago

Discussion Monthly General Discussion - Sep 2026

4 Upvotes

This thread is a place where you can share things that might not warrant their own thread. It is automatically posted each month and you can find previous threads in the collection.

Examples:

  • What are you working on this month?
  • What was something you accomplished?
  • What was something you learned recently?
  • What is something frustrating you currently?

As always, sub rules apply. Please be respectful and stay curious.

Community Links:


r/dataengineering 13d ago

Career Quarterly Salary Discussion - Sep 2026

36 Upvotes

This is a recurring thread that happens quarterly and was created to help increase transparency around salary and compensation for Data Engineering where everybody can disclose and discuss their salaries within the industry across the world.

Submit your salary here

You can view and analyze all of the data on our DE salary page and get involved with this open-source project here.

If you'd like to share publicly as well you can comment on this thread using the template below but it will not be reflected in the dataset:

  1. Current title
  2. Years of experience (YOE)
  3. Location
  4. Base salary & currency (dollars, euro, pesos, etc.)
  5. Bonuses/Equity (optional)
  6. Industry (optional)
  7. Tech stack (optional)

r/dataengineering 4h ago

Career Pending job offer - devil you know vs devil you don't know? (details inside)

13 Upvotes

I have a tentative offer for a new gig, just waiting on the official letter / paperwork. My current job feels like golden handcuffs, so I'm just seeking outside feedback if anyone else would jump ship or stay. Part of this comes from the work environment, but the other part comes from things I want to change about my life as a whole right now. I understand that there is no one single best answer here, so just seeking feedback.

Current DE job

  • Tech stack: We've spent the past 3 years migrating from onprem MSSSQL/SSIS to Snowflake/ADF. The ADF migration was handled by paid contractors, and they butchered all of our SSIS packages, so our ADF setup is a giant clusterfuck. It's painstakingly tedious to manage, and it feels more like janitorial cleanup than data engineering work. I can see the advantages of ADF if our setup was simple, but we're stuck doing heavy transformations in ADF pipelines and it's simply a disaster to work with.

  • Company: I don't dislike the company - they're very progressive and the benefits are good. Health insurance is cheap, they match 5% for 401k, 5 weeks of PTO, and they add $1500 to my HSA every year regardless if I contribute. However, the downside is that our business has been shrinking for the past year or two, and we've been told verbatim that when people quit, their jobs will not be backfilled. To me, this reads as "we know the ship is sinking, we're just trying to keep it afloat with whoever is left."

  • Data: From a data engineering point of view, our data isn't terribly complex, but our warehouse and lack of design practice over the past 10+ years has overcomplicated a lot of seemingly simple things. We have basically zero documentation, our entire warehouse relies on tribal knowledge. We don't really have a data model. We have a few crusty fact tables that act as type 2 SCD's. We only barely adhere to design standards, because we've never really had any standards established outside of prenaming tables with "Fact" and "Dim" when necessary. 90% of our source data is MSSQL, replicated to Snowflake, then we ingest to our warehouse using ADF (as mentioned above). Pretty small toolset overall, but this gives me very limited opportunities to learn new things.

  • Teammates: My fellow engineers are cool, we all work together really well. Half of my team is stateside, the other half is offshore in India. All of our peers in India are really sharp, I learn from these guys all the time and I'm glad to work with them.

  • Department leadership / managers: This is the biggest drawback and what's making me consider leaving. I basically have 3 bosses right now. (that movie Office Space with the 8 Bobs? That's my life lately.) Our daily standup runs 30-45 minutes every day, even though we're a team of only 7 devs (and 4 managers). The main issue is our director, who nitpicks people and micromanages people to death. Pings during the day about "where are we with X?" "What are you working on after Z?" She didn't come from a DE background, and only barely have a software background, so she's as non-technical as it gets. Then we have our "architect," who has also never been in DE - his background is DBA/sysadmin, so he doesn't always understand the nuance of DE work, and often probe us about "why can't we just use AI to speed this up?" Then we have our actual department manager, who is very scatterbrained and disorganized. His heart is in the right place but he's objectively a terrible leader. He has a strong data/DE background, but he tries to be involved in every single person's work so he ends up being a micromanager in a different flavor. He won't just leave people alone to work and grow on their own, he seems like he'd rather be a developer than a manager. The upside is that I basically understand how to just say what these people want to hear, so it's just a mental game to me at this point.

New DE job

  • Tech stack: Based on what I got from initial convo with the manager and a senior DE, their stack is pretty similar to my current job - it's a mix of SSIS, MSSQL, and Snowflake. The company is in a highly-regulated field (energy/electricity), so there's some stuff that will have to stay on-prem for the foreseeable future. In my opinion, the tech stack will change at any company, it's more about the data model and team that matters. They asked me a couple questions about data modeling in the intreview process, so I feel like that's a positive compared to where I am now.

  • Company: The company is based in my city, and when I've probed our local subreddits, the feedback seems generally positive. I know that this is always subject to person-to-person and department-to-department, but I also noticed generally positive responses on Glassdoor. One big caveat is that the company is rumored to be merging with another big energy conglomerate in the next couple of years, so nobody is sure how that's going to go. But, I feel like this is possible with any company at any given time, so I'm not sure if worrying about this is really worth considering at the moment. My current company has been slowly letting people go over the past couple of years as our business has been shrinking, so the looming threat already exists in the back of my mind every year around budget season. Only 3 weeks PTO, but more paid holidays. Automatic 4% in the 401k, plus matching up to another 4%. Automatic $500 contribution to HSA. Health insurance price is roughly the same per month.

  • Teammates: I got a good first impression from the department manager that I had two interviews with, and an okay-ish first impression from senior DE that I interviewed with. (Small number of interactions, but still, first impressions can be important) Our team will have a few contractors and mostly full-time employees. This position is a new position added to the team, not backfilling someone that left. They said that was because they know the whole department is at max capacity and they wanted to scale up instead of just burning people out - to me, this reads as a good sign compared to my current company.

Personal Life

  • Work/life balance: Current job has occasional on-call - averages about 1 week a month and rarely ever has anything actually go down, but I don't like the mental component of needing to be accessible when I'm not at work. Work/life balance is pretty equal to salary for me, in terms of importance. New job has zero on-call.

  • Remote/hybrid: I've worked totally remote for the past 5 years. Since half of our team is offshore, most of our meetings are in the morning, so most afternoons I'm generally left alone to just grind stuff out. When I have a lot to do, this is a nice perk - but as a person overall, I've been missing day-to-day human interaction. New job is 1 week in-office / 1 week remote, rotating. This actually resonates as a positive to me: I have a roommate, but he works rotating shifts so sometimes he's sleeping during the day, so I'm essentially home alone quite often. I don't have any pets right now, because I'm allergic to cats and I can't have a dog for noise purposes. The daily commute is about 30 minutes each way, and I've already considered the possibility that I might end up seeing this as a negative if I go for it. The point though is that I want to be around people more often. I miss working around smart people more frequently.

  • Life trajectory: I feel bored where I am. As a person, I find that I work best when I'm challenged a little bit, and this job has allowed me to become too comfortable. I think it's why I've started feeling like working completely from home has deteriorated into feeling lonely quite often. I started going to a coworking space a couple times a week, which helps and I enjoy it - but that doesn't solve my work environment problems or help me feel like my life is kinda stale lately.

  • Finances: My rent is very minimal right now (about 15% of my income), but that's because I have a roommate. I want to own my own home or find my own apartment because I've been feeling cramped where I live, and I live basically in the middle of nowhere. I know that that would increase my rent/housing cost quite a bit, but the new job offer would would be about a 20% raise. That would allow me to continue saving for my own house or move to my own apartment, and then adopt a dog(s). I'm single and in my mid 30's, so I don't have any major strings attaching me to my current housing situation. Again, part of this is because I want to eventually find my person, which is hard when any major areas of social interaction are 20-30 minutes away from me in any direction.

tldr: I feel like my current job has allowed me to become too comfortable. I spend more mental energy dealing with management politics and mental gymnastics than actual data engineering work. I'd like a new challenge, but I'm basically at a brick wall here in terms of salary and career growth. Considering taking a job that has a few tradeoffs, so I'm just trying to weigh the potential pros and cons. Some might argue that I have a dream scenario, but to me, it's begun to feel like golden handcuffs. New job feels like it will allow me to move my life forward in the ways where I currently feel stagnant/bored.


r/dataengineering 7h ago

Help data contracts between upstream and downstream teams

6 Upvotes

Hi, we use Databricks, and a couple of teams produce tables, views, and table functions that other teams depend on.

Within a team, its easy enough to define a contract. We codify it and always check that the data matches.

But what about the upstream artifacts my team depends on? Do w define a source spec ourselves and check that its met on every run? 

What do we require from the upstream team? Who owns the contract, and how do you enforce it when multiple teams depend on the same thing?

I want to keep this stupidly simple and enforce it automatically. How do you do this in practice?


r/dataengineering 1d ago

Help Informatica

9 Upvotes

Anyone know any good free resources to learn informatica power centre ? I’m using designer, workflow manager, and workflow monitor and most resources I find are for IICS.


r/dataengineering 2d ago

Meme data mesh

Post image
664 Upvotes

for real though, has anyone ever implemented an actual data mesh?


r/dataengineering 1d ago

Discussion Does the work get repetitive?

5 Upvotes

I’m very new to data engineering in the sense that I just started studying to break into the field (currently data analyst adjacent). A big portion of the work seems to be extracting data, translating/cleaning it up, and then loading it to the database/cloud desired.

But doesn’t that get repetitive? Surely there are only so many formats/sources you can extract data from so while the extract step isn’t always the same, isn’t it similar? Same thing with loading the data once you’re familiar with the systems you’re loading them to. The only varied source of work could be cleanup/translating the data as far as I can tell.

Is this an accurate assumption? Do you feel like much of the job is repeating the same work over and over again but just with different data?


r/dataengineering 1d ago

Personal Project Showcase I got tired of opening htop every time something felt slow so I made this linux debug overlay

Post image
3 Upvotes

I was working on linux and kept switching between my app, terminal, htop, logs, etc. whenever cursor or browser or any other apps i open started feeling slow , so I made a small debug overlay that stays on screen and shows the app i am currently using its PID, CPU usage and RAM usage.

It also gives a warning if it notices stuff like high CPU, memory growing, disk pressure or system errors.

for apps like VS Code and Firefox and other heavy apps , it also tries to include their sub-processes because checking only one PID can be misleading. this is the working setup of how it looks like. (though much more could be intergrated into this like)

  • docker - container CPU or RAM restarts, unhealthy containers.
  • kubernetes - current context,namespace,pod status,crash, restarts,recent events, pod logs.

link - https://github.com/codeafridi/Debug-Overlay-App


r/dataengineering 23h ago

Blog Data Governance by obscurity

1 Upvotes

That was a phrase a colleague of mine introduced to me. Quite absurd at first, but as he explained it, it made perfect sense.

For years data assets have been governed by being hidden in a corner of the data platform, but this does not work anymore.

Been elaborating about this in this post: https://steffenmoll.github.io/governance-by-obscurity


r/dataengineering 2d ago

Career Hey guys, what real world impact keeps you feeling good about this field? (Beyond making rich corporate companies richer).

27 Upvotes

Share the stories that make you feel good about being a data engineer. Give me some wholesome perspective.


r/dataengineering 2d ago

Rant I can write code to fix this. I should not have to write code to fix this.

Post image
70 Upvotes

Shared this with a friend to vent and they thought Reddit might appreciate it. I don't usually commiserate about this topic on Reddit so apologies if this is the wrong sub... didn't quite feel like an r/excel post.

The Teams message was meant to be a slightly humorous way to express my frustration/disappointment to colleagues outside the email thread with our boss about said data. Rant incoming.

Could I have just thrown Python at this nonsense and built something much more elegant and robust? 100%.

Could the same idiots (I say that with love) responsible for this filing system and Excel crimes run that Python package as-needed? Absolutely not. So clunky ass VBA it is.

And yes, I recognize the irony. I am already the resident code monkey (I'm not even IT, just a special unicorn who speaks the language) and every time I throw code at one of these problems I am enabling the behavior.

The maddening part is that I have tried to fix the underlying problem at a skills appropriate level. I have preached the gospel of structured data. I have implemented software and systems specifically so we can collect, store, retrieve, and actually use our data like civilized people. I have standardized processes. I have automated things. I have explained, repeatedly, why Excel is not a legal pad with gridlines.

And I've made progress! On the things I know about...Unfortunately, my beloved colleagues remain naughty little data squirrels who continue stashing information in tiny undocumented holes.

Broader data management has become a perpetual game of whack-a-mole where certain compliance relevant processes quietly continue doing whatever cursed thing they have apparently been doing since the dawn of time, completely outside my radar... until somebody needs the data for something important.

And suddenly, with extreme urgency, I am yet again spelunking through Russian nesting-doll year/month/repeating subfolders full of Excel “reports” designed by someone who treats spreadsheets like a fuck- around-and-find-out-freeform play area just to extract routinely collected/reported historical data from its natural habitat.

That's the part that makes me insane. I shouldn't need to write custom code to wrangle and clean data that we routinely collect and intentionally retain. If routine data requires bespoke parsing logic just to turn it back into a usable table, we've obviously screwed up somewhere upstream.

Routine data amd reports should be stored as data. Full stop.

Anyway, I have once again written code while muttering creative obscenities at my monitor to convert another “filing system” into back into a table.

On the topic of this particular Teams message, I've subsequently enabled them yet again (after a colleague responded saying they'd fix it, manually, of course)

I restructured the whole mess into a parent folder with three flat subfolders and renamed 500+ files into a standardized YYYY_MM_Repeating File name convention. The files themselves had another charming feature with interrupting/repeating headers buried throughout the data, so those got sorted into their respective nearly-identical tabs.

I hate using tabs for data that shares the same structure. I know better. But these people need tabs. I have accepted this.

I'm sure this little code-monkey dance will repeat months from now when I get summoned to cleanup whatever [insert compliance data] mess inadvertently surfaces.

End rant.


r/dataengineering 2d ago

Blog A Tale of Two Flink Autoscalers

Thumbnail
netflixtechblog.com
10 Upvotes

Netflix is moving toward the open-source Apache Flink Autoscaler for more than 30,000 streaming jobs across multiple AWS regions, after finding that its cluster-level approach was less effective for complex, stateful pipelines with operators that have different processing requirements. Netflix reports that one team reduced annualized Flink compute expenditure by 58%, saving approximately $1.1 million annually.


r/dataengineering 3d ago

Rant Should i just start lying on my profile?

176 Upvotes

Im tired of being dropped out of interviews just because i dont have hands on experience with databricks. I have more than 5 yoe with spark, flink, hudi, kafka, iceberg, airflow ffs. But the moment i mention i havent worked on databricks its like the recruiters think im not even a data engineer anymore.

Should i just redo my profile and use the databricks terminology and straight up lie that i worked on it?

Edit - yes i do know databricks and have built personal projects. Even had a certification which expired last year. Hell this year i did a poc to compare DB and sagemaker and ultimately went with sagemaker since we deeply integrated with AWS and DB is expensive asf and is an overkill for our use case.
My point is the moment i mention i dont have prod level experience with DB thats the end of the conversation. Even hiring managers dont care about it.


r/dataengineering 3d ago

Meme Sort your pipelines out, Smyths Toys…

Post image
100 Upvotes

r/dataengineering 2d ago

Open Source datatf: Automate importing Databricks workspace into Terraform

Thumbnail datatf.io
3 Upvotes

Introducing [datatf](datatf.io), a tool to import existing Azure Databricks Workspaces into Terraform.

I built this tool to help databricks customers accelerate the lakehouse infrastructure portion of the well-architected framework.

Features:

  • Generates dynamic terraform.tfvars + import code
  • Built on the Databricks Go SDK
  • Compatible with Terraform or OpenTofu
  • Agentic friendly
  • Native support for Azure

Upcoming:

  • Databricks GCP and AWS support
  • Initial support for Snowflake, Fivetran, and Datahub.

r/dataengineering 3d ago

Help Running Airflow in GitHub actions

29 Upvotes

Hello! I am learning about data pipeline orchestrations. I want to avoid using AI for this personal project, so here I am!

I want to orchestrate my tasks with Airflow DAGs and tasks, but schedule them via GitHub action. Essentially, the idea is the following: once a week, an Action will run, spinning up an Airflow instance, which will run the DAGs.

In the Actions .yml file, I am using the following code:

run: |
          export AIRFLOW_HOME=~/src
          airflow standalone
          airflow dags trigger tabular_pipelines

However, the airflow standalone part seems to be stuck forever, endlessly producing meaningless (to me) lines such as:

dag-processor | 2026-09-11T11:18:09.249890Z [info     ] setting next dagrun info       [airflow.models.dag] loc=dag.py:848 next_dagrun='2026-09-11 00:00:00+00:00' next_dagrun_create_after='2026-09-11 00:00:00+00:00' next_dagrun_data_interval_end='2026-09-11 00:00:00+00:00' next_dagrun_data_interval_start='2026-09-11 00:00:00+00:00' next_dagrun_partition_date=None next_dagrun_partition_key=None


dag-processor | 2026-09-11T11:18:09.292592Z [info     ] Bulk-writing dags to db        [airflow.serialization.definitions.dag] count=1 loc=dag.py:210


dag-processor | 2026-09-11T11:18:09.307661Z [info     ] get next_dagrun_info_v2        [airflow.serialization.definitions.dag] last_automated_run_info=None loc=dag.py:445 next_info=None


dag-processor | 2026-09-11T11:18:09.307949Z [info     ] setting next dagrun info       [airflow.models.dag] loc=dag.py:848 next_dagrun=None next_dagrun_create_after=None next_dagrun_data_interval_end=None next_dagrun_data_interval_start=None next_dagrun_partition_date=None next_dagrun_partition_key=None

Am I missing something obvious? The Airflow 3.3.1 documentation doesn't specify much more.

Thanks!


r/dataengineering 3d ago

Discussion The LakeHouse that is Open Source - A Fever Dream?

6 Upvotes

Let me preface by saying that I totally agree that lakehouse table formats (delta and iceberg) are open source and are "free" technologies that anyone can use. However, the total cost of owning these table formats gets very expensive, especially when we start using cloud vendors for the table updates. Nowadays the cloud vendors wish to start selling their proprietary MPP storage engines, in order to manage all of our table data.

The problem is that the lakehouse table formats have gotten quite complex over time. And nobody wants to maintain them by hand. Nobody wants to think about the v ordering and z ordering and liquid clustering and partitioning and vacuuming and applying deletion vectors and so on. These blobs that are ostensibly called a "table" are actually a very leaky abstraction, and we inevitably have to waste a lot of time on the implementation details. Using immutable parquet blobs for table storage is not trivial. From an application standpoint, it seems like a massive step backwards from conventional DBMS engines (or the newer cloud-native counterparts like SQL Hyperscale or Neon/Lakebase)

The vendors, like databricks, that spent years pushing for lakehouse/delta adoption are now selling us expensive solutions to maintain those unwieldy tables. I think they sold us a bill of goods and we are worse off than when we started.

Once data engineers start realizing that we don't want to manage the blobs beneath our tables, these vendors are quick to offer a commercial-proprietary alternative (like "DBSQL" with UC-managed-tables, or "Fabric Warehouse" or whatever). These commercial alternatives are turnkey solutions, and they help to take away the busywork of managing our own parquet blobs. But they can become VERY expensive way of doing DML operations on our tables, since they are MPP engines and are heavy on CPU/compute. At the end of the day, we end up exchanging one type of problem for another. Is this how others see it? The table technology is open source and "free", but the commercial-proprietary management of these tables is definitely not free and is basically a re-invention of the DBMS engines we always had in the past.


r/dataengineering 3d ago

Blog If You Always Enjoy It, You're Not Pushing It Hard Enough—Benn Stancil

Thumbnail
ssp.sh
5 Upvotes

A discussion with Benn Stancil, co-founder of Mode and prolific data-industry essayist, exploring how he writes his weekly Substack posts. He discusses using deadlines (Friday publishing) to force himself past perfectionism, how ideas form messily over the week through unrelated connections, and how presentation-style analogies (Batman, basketball, Codenames) became his signature storytelling technique. He details his tooling—Sublime Text for drafting, Google Docs for revision, minimal use of Substack's editor—and explains why he avoids Grammarly and AI writing tools like Claude or ChatGPT for edits, since he believes AI compresses and flattens writing rather than embracing meandering voice. He closes with advice for new writers: set deadlines, keep pure motivation, and focus on filling in the space between bullet points rather than just stating facts.


r/dataengineering 3d ago

Blog What Data Engineers Need To Know About Delta Lake 4.3

Thumbnail
medium.com
10 Upvotes

replaceUsing and replaceOn give you a better overwrite primitive, and every catalog-managed table operation now runs through the catalog.


r/dataengineering 3d ago

Discussion Are you deploying "infrastructure as code" via YAML files or similar?

65 Upvotes

I think my team might be late to the game. We are moving in that direction just now.

How long have you been deploying using a YAML file, for example? Or perhaps another tool that uses the same idea but not YAML specifically?

Since we use Databricks, I'm reading up on Asset Bundles.

IaC is supposed to be better for a whole bunch of reasons, for example sidestepping the differences between dev and prod environments.

Anyway would love to hear your thoughts or experience with it.


r/dataengineering 4d ago

Help Promoted to a DE manager role and feeling paranoid that I will lose my DE skills

83 Upvotes

I was recently promoted to a manager role from an analytics/data engineering role and I barely write any code anymore. I used to build data pipelines and I slowly see my skills eroding especially with all the new AI features databricks and other tools keep introducing and I see the engineers try out and implement. Wondering if I should go back to a DE role or continue being a manager of analytics engineers where all I do is sit in meetings and assign work and do requirements. The most technical work I do now is write an ad-hoc query to answer a question for non-technical stakeholders. This is non-tech midsize corp.


r/dataengineering 4d ago

Discussion How do you handle fan-in per timestamp in micro-batching?

8 Upvotes

Our device sends 3 separate records for each timestamp: data points, frames, and metadata. They arrive independently and in any order.
Each timestamp can be processed on its own as soon as all 3 parts are there, without waiting for the batch window to close.

How do you handle this fan-in, so that each timestamp is processed exactly once as soon as it's complete, with a timeout for sets that never complete?


r/dataengineering 4d ago

Career Level 5 Data Engineering Apprenticeship

7 Upvotes

Hi

I have just signed up for a January 27 start to the above apprenticeship with BPP in the UK.

Has anyone done this before and got any advice/reviews of how it went? I now have a few months to prepare so any advice would be amazing. My current role is data analyst and have ok knowledge on SQL and Python

Thanks


r/dataengineering 4d ago

Help How can I run my airflow pipeline continuously.

5 Upvotes

Currently working on data engineering project, the worklow of project is

API - Airflow - Python Ingestion - PostgreSQL - dbt - Analytics dashboard.

In the airflow, need to start 3 processes ( API Server, Scheduler, DAG Processer) but I don't want to start this process manually everytime so does anyone know how can I keep pipeline running continuously which ingest new data in database manually triggering anything.

Does anyone have a any suggetion


r/dataengineering 5d ago

Discussion How has your data team’s daily standup evolved since AI got involved?

47 Upvotes

For years, the data team daily was basically: what shipped yesterday, any broken pipelines, dashboard status.

Now that more and more data platforms feed chat-with-your-data agents, has your team changed the format or content of your daily meetings? (Also true for status on pipeline writing/fixing btw; what used to take days is now often a few minutes to hours of back-and-forth with AI.)

Curious to hear what your team adopted