CHOOSING & WORKING WITH AN AGENCY · September 2026 · ~11 min read
Setting expectations for the first 90 days
The first thirty days are access, measurement, and fixing what is plainly broken. Days thirty to sixty are production. By day ninety you should have movement on your easier searches, a working measurement setup, and a clear answer about what drives this business. You should not expect your hardest competitive term to have moved.
On this page
- 01What actually happens in the first thirty days?
- 02What should the review target for the quarter be?
- 03What should I see by day sixty?
- 04Should anyone be A/B testing my site in the first quarter?
- 05What is a fair test at day ninety?
- 06Which things move quickly and which do not?
- 07What to do this week
- 08When you do not need this
- 09Sources
- 10Related reading
- 11Questions about what your first ninety days should look like?
Almost every failed engagement I have heard described was a timeline disagreement that nobody wrote down. The client expected the phone to change by month two. The agency expected to be judged at month nine. Neither said so.
Write the ninety day picture down before the work starts. It costs one meeting and it removes the argument entirely.
01What actually happens in the first thirty days?
Three things, and only one of them is visible.
Access and inventory. Someone collects owner level access to your profile, analytics, Search Console, ads, and site, and lists what exists. This is unglamorous and it is where hidden problems surface, such as a domain nobody can log into or tracking that has been broken for a year.
The baseline. Current calls, forms, bookings, map impressions, and positions across a grid rather than a single number. This is the most valuable hour of the entire engagement, because everything after it becomes measurable.
Then the obvious repairs. Hours corrected, categories reviewed, missing services added, a broken form fixed, duplicate listings resolved. Category selection is the piece that pays back fastest, since your primary category does more for local ranking than almost anything else you can change. Darren Shaw's study of 1.8 million Google Business Profiles across 4,209 categories found the exact match category outranking every adjacent one, with the pattern repeating for almost every category in the data. It is a research task of an hour or two and a change that takes minutes.
If your first month feels slow and administrative, that is what it is supposed to feel like.
02What should the review target for the quarter be?
Worked example
A number, derived from your competitors rather than from a vendor's benchmark. This is the clearest piece of ninety day arithmetic available, so run it in week one.
Step one, count the competitors. Take the three businesses ranking above you for your main search and count how many new reviews each received in the last twelve months. Say the answer averages 24, which is two a month.
Step two, set the target at their rate plus one. That is Darren Shaw's published rule of thumb, and it gives you three a month, so nine across a ninety day engagement.
Step three, decide what nine reviews implies about your asking system. Shaw's other published guidance is that a business asking every customer should expect roughly thirty positive reviews for every negative one. At nine reviews in a quarter you are not yet at the volume where a single unhappy customer changes your average much, which is the honest reason not to panic about the first negative one.
Step four, know why the target is velocity rather than total. Review recency ranked eleventh and a sustained flow of reviews rather than bursts ranked fourteenth in Whitespark's 2026 survey of forty seven local search practitioners. Shaw's phrasing is that the moment you stop getting new reviews, local rankings start to slip, and that a negative review is better than no new review at all. Sterling Sky documented the same pattern from the other direction, in a case where a client's rankings dropped and the review flow had flat lined because the owner stopped rewarding staff for asking.
Where it breaks down: this is expert opinion and case observation, not a controlled test, and the survey's own author says the panel does not have access to the algorithm. Treat the target as a defensible operating goal, not as a lever with a known output.
What should I see by day sixty?
Production, and the first honest data.
Pages being published or rewritten on a schedule you agreed. Profile activity running. Review requests going out through whatever system your business already uses. Whatever technical items came out of the audit, being closed in a documented order.
You should also be getting your first real report against the baseline. Early numbers are noisy and should be presented as such. A report at day sixty that claims a clear result is overreading the data.
What you should feel by day sixty is the working rhythm. Whether requests get answered, whether the reports arrive on the stated date, whether the person doing the work can explain a decision plainly. That is more predictive of the next year than any metric at this stage, and it depends heavily on the contact cadence you agreed at the start.
04Should anyone be A/B testing my site in the first quarter?
Almost certainly not, and the rule everyone quotes for this is worth taking apart because the real reason is better than the folklore.
You will hear that A/B testing needs ten thousand monthly visitors. That number is not a research finding. It is traceable: it appears as a citation in Kohavi, Deng, Longbotham and Xu's KDD 2014 paper on running web experiments, and the reference resolves to a blog post by Neil Patel on QuickSprout, published on 14 January 2013. The four authors, from Microsoft and LinkedIn, immediately qualify it by saying the guidance should be refined to the metrics of interest. A blog post became an industry threshold by repetition.
The real rule is that sample size depends on the variance of the metric and the size of the change you want to detect. The same paper publishes a worked table from Bing that shows how far apart the answers can be for the same product. Revenue per user, a very skewed metric, needed about 114,000 users per variant to detect a 4.4% change. The capped version of the same metric needed about 9,700 users per variant to detect a 10.5% change. Sessions per user needed about 4,700 per variant to detect a 5.4% change.
Now run it against a local business. Say your site gets 600 visits a month and converts at 3%, which is 18 leads. Split into two variants, that is 300 visits and nine leads per variant per month. To reach 4,700 visits per variant, the least demanding line in that table, would take about 16 months (derived), by which point your season, your prices and your offer have all changed.
Say the domain gap out loud, because it cuts both ways. That table is search engine data at web scale, and a local lead form is a different metric with different variance. The transferable part is the method, not the numbers: ask for the minimum detectable effect and the variance before agreeing to a test, and if nobody can supply them, you are not running an experiment, you are changing the page and hoping.
What honestly replaces it in ninety days: qualitative review, heuristic fixes, and before and after measurement over long windows, described accurately as a comparison rather than a controlled test.
05What is a fair test at day ninety?
Four questions, and none of them is about a single ranking position.
Did the agreed deliverables get produced, counted against the scope. This one is factual.
Has anything moved on your easier searches. Less competitive terms, longer phrases, and positions in the parts of your service area where you were already close. Movement usually appears at the edges first.
Can you now see what happens. Calls tracked, forms attributed, the map grid comparable to the starting grid.
And do you know more about your business than you did. Which services people search for, which neighborhoods respond, what your competitors are doing that you were not.
Measure all four against your own trailing ninety days, never against a published average. There is no current benchmark for Business Profile performance worth the risk. The only large sample study of profile metrics is BrightLocal's, published in July 2019 on data from 45,264 businesses gathered between September 2017 and December 2018, and it has never been refreshed. Google has since retired several of the metrics it was built on, and views are now counted as deduplicated unique visitors rather than raw views, so its conversion figure cannot be reproduced on a modern account. Your own starting numbers are the only valid comparison you have.
A ninety day engagement that produces no ranking change but a working measurement system has still earned its money. That statement is also extremely convenient for anyone selling ninety day engagements, this one included, so here is the test that keeps it honest: the measurement setup has to be something you can log into and read without the agency present, and it has to survive their departure. If it does not, it was a report, not a system.
What you should not expect: your most contested head term in your densest market, a revenue line you can trace cleanly, or stability in AI answers. That last one confuses people, and it should not, because AI results change between runs of the same question by design. The St. Gallen measurement study found engine to engine differences that are large in themselves, with Gemini the most stable of the four tested and ChatGPT the least, and it found that ChatGPT activated web search on only 42.2% of runs. Checking daily produces anxiety, not information.
06Which things move quickly and which do not?
Fast, within weeks. Profile completeness, categories, services and products, hours, photos, and review flow once the asking system is running. Technical fixes on your own site. Anything that was simply missing.
Medium, one to three months. Positions on longer, less competitive searches. Pages that answer a specific question that nobody local had answered. Conversion improvements on pages that already get traffic.
Slow, six months or longer. The most competitive term in a dense market. Authority built through mentions and links. Anything requiring a body of content rather than a page.
The uncomfortable one: some things do not move because of search work at all. If your listing looks fine and your calls go unanswered at 5pm, the constraint is the phone. A good agency will tell you that in the first month rather than sell you another quarter.
07What to do this week
Before the engagement starts, ask for the ninety day picture in writing. One page, three sections, thirty sixty ninety, with what gets produced in each and what is expected to have changed by the end.
Run the review arithmetic yourself. Count your top three competitors' new reviews over the last year, divide by twelve, add one, multiply by three. That is your quarter's target and it took four minutes.
Then ask the question that should be uncomfortable for whoever is selling you the quarter, and that includes us. What specifically would count as failure at day ninety, written as a number, and what happens to the fee if it is missed? SEOD does not tie its fee to a ranking outcome, because nobody honestly can, and that means we have to answer the first half of the question in writing or admit the checkpoint is decorative. An agency that cannot imagine failure has not thought about accountability.
Then agree what you owe. Approvals within a set number of days, photos, access, and a monthly half hour. Write your side down too, because half of stalled engagements stall on the client side, and an honest look at the cost of doing it yourself versus hiring starts with the same accounting of your hours.
If you have not yet chosen, the questions to ask before hiring belong in the same conversation as the timeline.
Be honest with yourself
When you do not need this
If you are buying a defined project rather than an ongoing engagement, judge it on the deliverable and the completion date. Ninety day frameworks do not apply to a two week job.
If you are in a genuinely uncompetitive category and were simply invisible, the first thirty days may produce most of the result you were looking for. Do not let anyone talk you into a year of work when the fix was a category and a services list.
And if you have run this play before with a previous vendor and already have a clean baseline and working tracking, skip ahead. You are starting at day thirty, and you should pay for that accordingly.
Sources
- Whitespark, "Review Recency is the Most Underrated Local Ranking Factor in 2025," Darren Shaw, 2 May 2025. Source of the competitor rate plus one rule, the thirty to one ratio, and the statement about rankings slipping when reviews stop. Vendor research and practitioner opinion.
- Sterling Sky, Joy Hawkins, "Does Review Recency Impact Ranking?". The case where a client's review flow flat lined and rankings followed. Practitioner case study, small sample.
- Kohavi, Deng, Longbotham and Xu, "Seven Rules of Thumb for Web Site Experimenters," KDD 2014. Peer reviewed. Source of the Bing sample size table and of the citation that traces the ten thousand visitor rule to a 2013 blog post. Web scale desktop data, not local business data.
- BrightLocal, "Google My Business Insights Study," July 2019. 45,264 businesses, data from September 2017 to December 2018. Cited here as the benchmark that has aged out. Vendor research: BrightLocal sells local SEO software.
- Darren Shaw, "What 1.8 million Google Business Profiles tell us about local SEO success," Search Engine Land, 25 August 2026. 1.8 million profiles, 4,209 categories. Vendor research with a very large sample.
Related reading
- What a marketing retainer should actually include. The scope document that should sit underneath this timeline.
- Signs your current agency has stopped doing the work. The checks to run once the first ninety days are behind you.
- What a good discovery process looks like. What should be happening during the slow administrative first month.
- How to read an agency's case studies critically. Why the ninety day stories in a pitch deck rarely include the starting numbers.
Questions about what your first ninety days should look like?
Email me at eric@seod.com with your start date and the scope you have been given, and I will send back a checkpoint list for day thirty, day sixty, and day ninety that you can hold your agency to. It is written for you to use with them, not with me. No pitch attached and nothing sent afterward.
There is more on hiring and working with agencies alongside it.