SEO Agency Selection: RFP Templates and Evaluation Criteria
On this page
- Pre-RFP work shapes the outcome
- The RFP document structure
- Start from Google’s own questions
- Questions that separate strong from weak answers
- Weighted scoring across several evaluators
- Reference checks that check
- Red flags, and the contract terms that matter
- Frequently asked questions
- Related posts:
Excellent and incompetent SEO agencies can look identical in a pitch. Both can bring polished decks, confident accounts of past wins and someone who can read the room. The defense is a selection process built to surface capability a good pitch finds hard to fake: scenario questions that reward analytical thinking, weighted scoring by several independent evaluators, reference checks that go past the curated list, and a clear view of the red flags that separate operators from salespeople. If your process is “watch three pitches and pick the one you liked,” you have optimized for charisma.
Much of the work that shapes the outcome happens before any agency sees your RFP, and the questions that produce signal are hard to write but easy to score.
Pre-RFP work shapes the outcome
A bad engagement can be lost before the RFP goes out, because the buyer hasn’t agreed internally on what they are buying. Settle five things first:
- Objectives. What business result defines success: revenue-attributable organic traffic, rankings for commercial pages, lead volume, defending market share? “More traffic” isn’t an objective; it is a failure mode that can produce traffic with no revenue behind it.
- Budget range. Decided internally before proposals arrive, so you judge fit instead of being anchored by the first quote.
- Decision authority. Who scores, who signs, who can veto. Ambiguity here can let the most charismatic pitch win by default.
- Deal-breakers. Non-negotiables, such as industry experience, in-house or offshore execution, and reporting cadence, stated up front.
- Current-state documentation. Your traffic, rankings, technical baseline and prior work, shared so agencies can propose against reality. An agency forced to propose blind has little to offer beyond boilerplate, and you lose the chance to judge its thinking.
The RFP document structure
A good RFP produces comparable, signal-rich responses:
- Overview and context: your business, market and why you’re engaging now.
- Objectives and success criteria: the outcomes, and how success will be measured.
- Scope in and out: what the agency will own and, as important, what it won’t.
- Requirements and constraints: platform, compliance, reporting, timeline, work with internal teams.
- Proposal instructions: format, length and the specific questions to answer, so responses are comparable rather than brochures.
- Evaluation criteria: how you’ll score. Transparency can attract serious agencies and filter out the ones that won’t invest in a real response.
Start from Google’s own questions
Google’s guidance on hiring an SEO lists questions to ask a provider. Put them in the RFP:
- Can you show me examples of your previous work and share some success stories?
- Do you follow the Google Search Essentials?
- What kind of results do you expect to see, and in what timeframe?
- How do you measure your success?
- What’s your experience in my industry, and in my country or city?
- How can I expect to communicate with you?
- Will you share with me all the changes you make to my site, and provide detailed information about your recommendations and the reasoning behind them?
They set a baseline. The questions that separate strong agencies from weak ones go further.
Questions that separate strong from weak answers
The aim is questions that are hard to answer well without real capability, so a weak agency has a harder time bluffing through them.
Strategy. “Given this limited snapshot of our site and market, which three opportunities would you prioritize first, and why?” A strong answer names specific, defensible hypotheses and reasons from the data you gave. A weak answer recites a generic methodology, or defers every judgment to a discovery phase.
Technical. “Walk us through how you would diagnose a JavaScript site whose content isn’t being indexed,” or “how you would approach a Core Web Vitals problem.” A strong answer matches Google’s own documentation. Google’s JavaScript SEO basics describes three main phases, crawling, rendering and indexing; pages can sit in the rendering queue for a few seconds or longer, and Google uses the rendered HTML to index the page. So a good diagnosis compares the raw and rendered HTML, for example with URL Inspection, before blaming anything else. On performance, a strong answer names the current metrics: Google’s Core Web Vitals documentation says to aim for LCP within the first 2.5 seconds, INP of less than 200 milliseconds and CLS of less than 0.1. Google’s Search updates log records that INP replaced FID as a Core Web Vital in March 2024, so an answer built around FID is out of date.
Process. “What do your first 90 days look like?” and “How do you handle a recommendation that conflicts with what our engineering team wants to build?” Strong answers describe a concrete sequence and a real account of working through organizational friction. Weak answers are vague timelines and conflict-avoiding non-answers.
Results. “Tell us about an engagement that failed or underperformed, and what you learned.” It is hard to fake. A strong agency gives a specific, honest account with a real lesson. An agency that claims never to have failed invites a closer look at its client list.
The pattern across all four: strong answers are analytical, specific and willing to admit uncertainty or failure; weak answers are generic, deflecting and allergic to admitting a miss.
Weighted scoring across several evaluators
Score with a rubric, not a gut feeling, and have three to five people score independently before comparing notes, so one charismatic pitch has less room to dominate. Weight the criteria by what predicts success in your situation. For example, you might weight strategic thinking and technical capability more heavily, with process, relevant results, fit and price each carrying less. The weights are yours to set; the discipline is that they are set in advance, applied consistently, and that each evaluator writes down the reason behind a score. Written reasons let you defend the decision later and catch the evaluator who scored on impressions.
Reference checks that check
Provided references were picked by the agency, so treat them as a floor, not the evidence.
- Ask the provided references hard questions: what went wrong and how the agency handled it, whether results held after the engagement, how the agency responded when an approach wasn’t working, and whether they would hire them again for a harder problem.
- Find clients who aren’t on the list, through case studies, LinkedIn or industry contacts, including clients who left. Why someone left an agency can tell you more than why someone stayed.
Red flags, and the contract terms that matter
Google’s hiring guidance is direct: “No one can guarantee a #1 ranking on Google,” and it warns about SEOs that claim to guarantee rankings, allege a “special relationship” with Google or advertise a “priority submit” to Google. Treat those as disqualifying. Add:
- Vague methodology the agency won’t explain (“proprietary techniques”).
- Opaque pricing with no line of sight into what you’re paying for.
- The senior team that pitched disappearing after signing, replaced by juniors.
- No failure they can name.
- Pressure tactics and artificial deadlines.
- A copy-paste proposal that doesn’t engage your situation.
Before signing, review with counsel the terms that matter more than the monthly fee: termination and notice, ownership of content and deliverables, any exclusivity or non-compete restrictions, and liability. A strong process improves the odds of a good outcome; it doesn’t guarantee one, and any agency promising that it does has just shown you a red flag.
Frequently asked questions
How many agencies should receive the RFP?
Enough to compare meaningfully and a small enough set that evaluators can read every response in full. Three to five is a set a small evaluation team can score properly.
Which question is especially hard to fake?
The failure question: “Tell us about an engagement that underperformed, and what you learned.” A specific, honest answer shows judgment; a claim of no failures is a warning.