[ETR #16] Better Data Solutions In 5 Steps


Extract. Transform. Read.

A newsletter from Pipeline

Hi past, present or future data professional!

When you apply to data analysis, data engineering or data science jobs, you likely consider factors like company name, culture and compensation. Caught up in the excitement of a fresh opportunity or compelling offer you’re neglecting an important part of your day-to-day reality in a new role: What stage of data maturity the organization is in. If you’re looking for experience building something new from the ground up, you likely won’t find it in a company that has a years-old established cloud infrastructure. If you’re inexperienced, you might also feel lost in a company that is still conceptualizing how it is going to establish and scale its data infra.

While I personally arrived at a team and organization in its mid-life stage, I’ve had opportunities to discuss, examine and advise those who are considering how they can make an impact at an earlier-stage company in both full-time and contract roles. This compelled me, after a transatlantic flight, to compile a framework you can use to conceptualize anything from an in-house data solution to full-fledged infrastructure.

Phase 1

Discovery - Extensive, purposeful requirements gathering to make sure you are providing a solution and, more importantly, a service, to an end user.

Phase 2

Design - You can’t begin a journey or a complex technical build without a road map; take time to make a wish list of must-have data sources and sketch your architecture before writing line 1 of code.

Phase 3

Ingestion - Build your pipelines according to best practices with a keen eye on cost and consumption; expect this to take 6-12 months depending on your work situation.

Phase 4

Downstream Build - Going hand-in-hand with requirements gathering, consider how your target audience will use what you’ve built; might it be better to simplify or aggregate data sources in something like a view?

Phase 5

Quality Assurance And Ongoing Tasks - Even though your pipelines and dashboards will be automated initially, nothing in data engineering is 100% automated. Components will break. You’ll be expected to fix them. And assure it doesn’t happen again.

These 5 phases aren’t meant to be strict rules for building data infra. But they should get you thinking about how to build something purposefully so you can spend your time dealing with angry code–not stakeholders.

Dive into the framework here.

Here are this week’s links:

Until next time–thanks for ingesting,

-Zach Quinn

Extract. Transform. Read.

Reaching 20k+ readers on Medium and over 3k learners by email, I draw on my 4 years of experience as a Senior Data Engineer to demystify data science, cloud and programming concepts while sharing job hunt strategies so you can land and excel in data-driven roles. Subscribe for 500 words of actionable advice every Thursday.

Read more from Extract. Transform. Read.

Hi fellow data professional!' Today I’m turning the newsletter over to my friend Ken Jee (writer of AI Survival Guide, creator of Newsletter Hero) to share how he cuts through the noise of shiny AI products to find tools that enhance technical work. My Simple Framework For Adopting AI Tools Ken Jee As new AI tools launch almost daily, a quiet tax is emerging. Decision fatigue. Every new model, agent, or workflow tool carries the same implicit question. Should I switch, or should I go deeper...

Hi fellow data professional! Quick question: How much could I pay you to switch your job? Conventional wisdom in the tech industry in the last handful of years is that the way to supercharge growth and max out your career earnings is to frequently change jobs. On average, job switchers could and should target an increase of 15-20% of their current salary. But in a rocky economy (at least here in the U.S.), career experts are urging would-be switchers to consider the benefits of a stable role...

Hi fellow data professional and Happy New Year! In the second half of 2025, I made a radical choice: I (largely) stopped blogging. Over the past year, Medium (where I host my content) made a series of changes that de-prioritizes technical content, leading to the departure of several major publications, including Toward Data Science. Pair that platform disillusionment with a bit of burnout, and the result is a feeling that it’s time for a change. For 75+ weeks, I’ve preferred concise,...