What CRO services actually cover
"Conversion rate optimization services" is a loose term. Two agencies can both sell it and be doing completely different work at completely different prices. Before you compare quotes, it helps to know what's actually in scope.
A real engagement usually includes some combination of:
- Measurement setup and validation. Confirming your analytics agree with your Shopify orders, that funnel events fire, and that the numbers you're about to make decisions on are trustworthy.
- Quantitative analysis. Conversion split by device, traffic source, landing page and funnel step, over a period long enough to be meaningful.
- Qualitative research. Session recordings, heatmaps, on-site search logs, customer service transcripts, and sometimes user testing.
- A prioritised backlog. Findings ranked by expected value, not by how easy they are or how much they annoy someone internally.
- Implementation. Theme and template changes, which may or may not be included — this is the single most common scope gap in CRO proposals.
- Testing. Running experiments, calling them properly, and documenting what was learned including the losses.
- Reporting. Against a baseline that was agreed before work started.
If a proposal covers research and recommendations but not implementation, you are buying a document. That can be the right purchase — but know that's what it is, and know who is going to build the changes.
How engagements are usually structured
Three shapes cover most of the market.
A one-off audit. A fixed-scope review producing a prioritised list of findings. Useful when you suspect there's a problem but don't know where, or when you need something concrete to get budget approved. The risk is that audits are easy to sell and easy to leave in a drawer. Ask what happens after delivery.
A retainer. Ongoing monthly work — research, testing, implementation, reporting. This is where most sustained gains come from, because CRO is iterative and a single pass rarely finds the biggest win. The risk is drift: retainers that quietly become maintenance. Agree what "a good month" looks like before you start.
A project. A defined piece of work with a defined end — a checkout rebuild, a product page redesign, a mobile overhaul. Clear scope, clear finish, but no mechanism for learning what happened afterwards unless you build measurement into it deliberately.
Many engagements start as an audit, prove the case, and convert to a retainer. That is a reasonable path as long as the audit is priced as an audit and not as a sales document you paid for.
What a good first three months looks like
Specifics matter more than promises. Roughly what you should expect:
- Month one is mostly measurement. Not glamorous, and the month clients most often want to skip. If your analytics and your orders disagree, every decision after this is guesswork. Expect a baseline document at the end of it, not a list of wins.
- Month two is research and the first changes. The obvious broken things get fixed rather than tested — a broken mobile variant selector doesn't need an experiment, it needs repairing. Genuine unknowns go into the testing queue.
- Month three is the first real tests reading out. Some will be inconclusive. That is normal and not a failure; an agency that reports only wins in month three is selecting what it shows you.
If someone promises a specific percentage lift in a specific timeframe before they have seen your analytics, they are guessing. There is no way to know the size of the opportunity before looking at where the drop-offs are.
How to judge whether it's working
Blended conversion rate is a bad primary scorecard, because it moves for reasons that have nothing to do with the work. A traffic-mix shift toward cold paid social will drag it down while the site gets better. A branded-search spike will lift it while nothing changes.
Better signals, in rough order of usefulness:
- Movement in the specific metric each change targeted. If the work was on the product page, look at add-to-cart rate on that template — not sitewide conversion.
- Conversion by device, tracked separately. Mobile and desktop are effectively different stores and improvements rarely land equally.
- Test velocity and test quality. How many experiments ran, how many reached significance, how many were called properly. A programme running two well-powered tests a month is healthier than one running eight underpowered ones.
- Documented learning. Including the losses. A year of CRO should leave you with a written body of knowledge about your own customers, not just a changed website.
- Revenue per session. Harder to game than conversion rate, and it catches average-order-value work that conversion rate alone misses.
Agree which of these you're measuring, and the baseline, before work starts. Retrofitting a baseline after the fact is how disagreements about results begin.
What you can do yourself
Being honest about this matters more than protecting scope.
Plenty is genuinely DIY. Splitting your conversion rate by device and looking at your own store on a real phone will usually surface something obvious within ten minutes. Pulling your top internal site searches and counting how many return nothing is free and frequently revealing. Auditing your app stack for things nobody uses generally improves speed more than any single optimisation. Our guide to Shopify CRO walks through the order to work in.
Where outside help earns its cost:
- You have enough traffic for real testing and no one internally who can design and call an experiment properly.
- The problems are in the theme or checkout and you have no development capacity.
- You have looked and cannot see it — which is common, because you are too close to your own store.
- You have run tests and can't tell whether the results mean anything.
Where it usually isn't worth it yet: stores without enough traffic for tests to reach significance. Below a few thousand sessions per variant, A/B testing produces confident nonsense. An agency that will sell you a testing programme at that volume is selling something that cannot work. The honest advice at that stage is to fix the obvious problems, do qualitative research, and spend the money on traffic.
Questions worth asking before you sign
- "Is implementation included, or do we build it?" The most common scope surprise.
- "What's your minimum traffic threshold for testing?" An agency without one hasn't thought about statistical power.
- "Who actually does the work?" Ask to meet them, not the salesperson.
- "What does a month look like when a test loses?" Reveals whether they've run real programmes.
- "Show me a case study where the result was ambiguous." Real work has attribution ambiguity in it.
- "How do you handle our checkout, and on what plan?" Checkout customisation depth is plan-dependent. If they don't ask what plan you're on, they don't know the platform.
- "What happens to the testing tooling if we stop working together?" You should keep your own data and your own experiment history.
- "Where did that benchmark come from?" If a pitch quotes an industry average conversion rate, ask for the source. Most circulating figures come from analytics vendors measuring their own self-selected customers, and Shopify does not publish an average for Shopify stores at all.
Red flags
- A guaranteed percentage lift, quoted before anyone has seen your analytics.
- A case study deck where every test won.
- Recommendations that could have been written without opening your store.
- Testing proposed on traffic that cannot support it.
- Reporting that only ever shows blended conversion rate.
- Pressure to install a heavy third-party testing script when Shopify's native experiments may cover the need.
What it costs, honestly
We're not going to publish a rate card in an article, because the honest answer is that CRO pricing varies more than almost any other agency service and quoted ranges from other sites are close to meaningless. What actually drives the number:
- Whether implementation is included. The single biggest factor. Research-only is a fraction of the cost of research-plus-build.
- Your traffic volume. More traffic means faster tests, more of them, and more analyst time.
- Your platform depth. A standard theme is different work from a headless build or heavy checkout customisation.
- Whether measurement needs fixing first. Frequently it does, and that's real work before any optimisation starts.
- Cadence. A quarterly review and a weekly testing programme are different products.
The useful comparison isn't cost against another agency's cost. It's cost against the revenue currently walking out of your checkout. A store doing 50,000 monthly sessions at a 2% conversion rate and a $50 average order is making $50,000 a month; moving that to 3% is $75,000 from the same traffic. Whether that's realistic for your store is exactly what the first month is meant to establish — but it's the arithmetic that tells you what the work is worth.
Tooling
The stack is less important than most vendors suggest, and it has changed recently in one significant way.
Shopify now has native A/B testing. Rollouts includes experiments that compare a treatment against a control across your main theme and your checkout and accounts pages. That removes the main historical reason for putting a third-party testing script in the render path, which was itself a common cause of the flicker and speed penalties that ate the gains. Shopify's documentation doesn't state plan availability on the main page, so check your admin.
Beyond that: Shopify Analytics and GA4 for quantitative data, a session-recording and heatmap tool for qualitative, and PageSpeed Insights for performance. Google Optimize was discontinued in 2023, so any guide still recommending it is out of date. Adding tools rarely fixes a CRO problem; the constraint is almost always attention, not software.
Common questions
What do Shopify CRO services include?
Typically measurement validation, quantitative and qualitative research, a prioritised backlog, testing, and reporting. Implementation is sometimes included and sometimes not — check explicitly, because it's the most common scope gap.
How long before we see results?
Simple fixes can land in weeks. A structured testing programme usually needs a few months, because tests need enough traffic and time to reach significance. Anyone promising a specific lift on a specific date before seeing your data is guessing.
What's a good conversion rate to aim for?
There isn't a reliable published benchmark for Shopify stores — Shopify doesn't publish one, and the figures in circulation come from analytics vendors measuring their own customers. Aim at your own baseline, split by device and traffic source.
Do I need enough traffic before CRO is worth it?
For A/B testing, yes — realistically thousands of sessions per variant. Below that, the useful work is fixing clear problems and doing qualitative research, not running experiments.
Can we do it in-house?
Much of it, yes, and the diagnostic work costs nothing but time. Outside help earns its keep when you need experiments designed and called properly, when the fixes are in the theme or checkout and you have no development capacity, or when you've looked and can't see it.
Working with us
First Pier is a Shopify and Shopify Plus agency in Portland, Maine. Our Shopify CRO service covers the full sequence above — measurement first, then research, then a prioritised backlog we implement and test.
The part worth knowing before you call: we spend the first month making sure your numbers are real, and we will tell you if we think your store doesn't have the traffic to justify a testing programme yet. If you'd like a view on where your store is losing people, get in touch.





.png)
.png)
