You're running a growing business. Sales come through Shopify, marketing runs on Mailchimp, customer support lives in Zendesk, and your accounts team works in Xero. Each system has its own dashboard, its own reports, its own version of the truth. When someone asks a simple question like "how much did we spend acquiring customers who've now churned?", you end up downloading three CSV files and spending an hour in Excel.
This is the problem ETL is meant to solve. But do you actually need it?
What ETL actually means
ETL stands for Extract, Transform, Load. It's the process of taking data from one place, changing it into a useful format, and putting it somewhere else.
Extract means pulling data out of a source system. That might be reading orders from your Shopify database, pulling email opens from Mailchimp's API, or downloading transaction records from your payment processor.
Transform means cleaning and reshaping that data. Perhaps you need to convert American date formats to British ones, merge customer records that appear in multiple systems, calculate derived metrics like customer lifetime value, or filter out test transactions.
Load means putting the processed data somewhere useful. Usually that's a data warehouse (a database designed for analysis rather than day-to-day operations), but it might also be a spreadsheet, a dashboard tool, or even another operational system.
A simple example: you extract yesterday's orders from Shopify as JSON, transform them into a clean table with columns for order date, customer email, total value, and product category, then load that table into a Google Sheet that feeds your morning sales dashboard.
When you probably don't need ETL
Most small businesses don't need formal ETL pipelines. We've seen plenty of companies waste time and money building data infrastructure they never properly use.
If your data volumes are small (a few hundred transactions per day), your questions are straightforward (weekly revenue, top products, customer counts), and you're comfortable with manual exports, then native reporting tools usually suffice. Shopify has perfectly adequate built-in reports. Google Analytics shows you website traffic. Xero produces financial statements.
The usual pattern we see: someone reads an article about "data-driven decision making", gets excited about dashboards, commissions an expensive data warehouse, and then... still makes decisions based on gut feel and the same three metrics they've always watched. The infrastructure sits unused because the business hasn't actually developed data-literate habits.
We worked with a small subscription box company last year who wanted a "proper data platform". They had about 500 active subscribers and were adding maybe 50 per month. When we asked what questions they couldn't currently answer, the answer was basically "none". Their existing tools showed them everything they needed. We talked them out of the project and saved them several thousand pounds.
When ETL starts making sense
ETL becomes valuable when you face specific practical problems that simpler tools can't solve.
First scenario: you need to combine data from multiple sources to answer business questions. A clothing retailer we work with wanted to understand marketing effectiveness. They needed to match Google Ads spend (from Google Ads) with website sessions (from Google Analytics) with actual purchases (from Shopify) with customer lifetime value (calculated from historical orders). No single tool could show this picture. An ETL pipeline pulls everything into BigQuery where they can join it together.
Second scenario: your data volume makes manual work impractical. If you're downloading and processing CSV files every morning, and that task takes 20 minutes, you're spending over 80 hours per year on data janitorial work. At that point, automating the process pays for itself quickly. One of our clients processes several thousand B2B orders per day across multiple sales channels. Manual reporting simply wasn't feasible anymore.
Third scenario: you need historical data that source systems don't retain. Many SaaS tools only keep detailed data for 90 days or a year. Mailchimp doesn't store your complete email history forever. If you want to analyse trends over multiple years, you need to be copying that data out regularly and storing it yourself. A law firm we work with needed to track case patterns over five years, but their case management system only retained 18 months of detailed records. ETL solved this by archiving everything to a SQL Server database.
Fourth scenario: you're spending serious money based on data, and you need confidence in the numbers. If you're allocating a six-figure annual marketing budget across channels, you can't afford discrepancies between platforms or rely on approximate figures. You need a single source of truth. ETL pipelines with proper validation and error handling give you that confidence.
What ETL actually looks like in practice
The simplest ETL pipeline might be a scheduled script. Every night at 2am, a Python script runs on a small server. It calls the Shopify API to fetch yesterday's orders, does some light processing (maybe it geocodes customer addresses or categorises products), and writes the results to a Google Sheet. Total cost: a few pounds per month for the server, perhaps a day of development time to build it.
For growing businesses, we typically build pipelines using dedicated ETL tools like Fivetran or Airbyte, which handle the "extract" part. These tools know how to connect to hundreds of common business systems and handle annoying details like pagination, rate limiting, and authentication. They cost money (usually starting around £50-100 per month), but they're far more reliable than maintaining custom API integrations yourself.
The "transform" step often happens using a tool called dbt (data build tool), which lets you write SQL queries that run in sequence to clean and reshape your data. This keeps all your business logic in version control and makes it reviewable and testable.
The "load" step goes into a data warehouse. For small businesses, we usually recommend Google BigQuery or a PostgreSQL database, depending on technical requirements and budget. BigQuery is particularly good because you only pay for what you use. If you're storing a few gigabytes and running occasional queries, you might spend £5-10 per month.
Then you connect a dashboard tool (Power BI, Looker Studio, Metabase) to the warehouse, and suddenly you have live reporting that updates automatically.
The realistic middle ground
Most growing businesses don't need full ETL infrastructure immediately, but they do need something better than manual CSV exports.
The approach we usually recommend: start with native integrations where they exist. Tools like Zapier or Make can connect systems together without code. If Shopify has a built-in Slack integration that posts your daily revenue, use that instead of building something custom.
When you outgrow simple integrations, consider managed ETL services before building custom pipelines. Fivetran, Stitch, and similar services aren't cheap, but they're far cheaper than hiring a data engineer or spending your own time maintaining scripts that break when APIs change.
Only invest in custom ETL infrastructure when you have specific requirements that off-the-shelf tools can't meet, or when the volume of data makes managed services prohibitively expensive.
Questions to ask yourself
Do you currently make decisions based on incomplete information because getting the complete picture is too much effort? That's a signal you might benefit from ETL.
Are you spending multiple hours per week manually combining data from different systems? That's time you could automate away.
Do you need to track trends over years, but your current tools don't retain historical data? That's a preservation problem ETL can solve.
Are you frequently discovering discrepancies in your data (two systems showing different numbers for the same metric) and you need a single source of truth? ETL can establish that.
On the other hand, if your main frustration is that your reports aren't pretty enough, or you feel like you "should" have better data infrastructure because other companies do, you probably don't need ETL yet. Spend your time and money elsewhere.
What to watch out for
The biggest mistake we see is building data infrastructure before developing data habits. ETL doesn't create a data-driven culture. It just makes data available. If your team doesn't currently use the reports they have access to, they won't suddenly start using reports just because they're automated and prettier.
Start by proving you'll actually use better data. Spend a month doing the manual work: download those CSVs, build those Excel pivots, answer those cross-system questions by hand. If you're doing this regularly and finding it valuable, then automate it. If you do it once and never look at the results again, you just saved yourself a lot of wasted effort.
The second mistake is underestimating maintenance. Data pipelines break. APIs change. Business logic evolves. Someone needs to own this infrastructure and keep it running. For most small businesses, that means either using very reliable managed services or working with an agency like us who can provide ongoing support. A custom-built pipeline that breaks and nobody knows how to fix is worse than no pipeline at all.
Be realistic about what you'll actually maintain. A simple nightly script that loads data into a spreadsheet and occasionally breaks is often more valuable than an elaborate pipeline that nobody understands well enough to fix when something goes wrong.