SEO Agency Selection: RFP Templates and Evaluation Criteria

On this page

Excellent and incompetent SEO agencies look identical in a pitch. Both have polished decks, confident accounts of past wins, and a closer who can read the room. The only reliable defense is to engineer the selection process to surface capability that a good pitch cannot fake: specific scenario questions that reward analytical thinking, weighted scoring across multiple independent evaluators, reference checks that go past the curated list, and disciplined recognition of the red flags that separate operators from salespeople. If your process is “watch three pitches and pick the one you liked,” you have optimized for charisma.

The work that determines the outcome happens before any agency sees your RFP, and the questions that produce signal are hard to write but easy to score. Design for both.

Pre-RFP work decides the outcome

Most bad agency engagements were lost before the RFP went out, because the buyer had not aligned internally on what they were buying. Settle five things first.

  • Objectives. What business result defines success: revenue-attributable organic, commercial-page rankings, lead volume, market-share defense? “More traffic” is not an objective; it is the failure mode that produces traffic with no revenue behind it.
  • Budget range. Decided internally before proposals arrive, so you evaluate fit rather than getting anchored by whoever quotes first.
  • Decision authority. Who scores, who signs, who can veto. Ambiguity here lets the most charismatic pitch win by default.
  • Deal-breakers. Non-negotiables (industry experience, in-house versus offshore execution, reporting cadence) stated up front, not discovered mid-engagement.
  • Current-state documentation. Your traffic, rankings, technical baseline, and prior work, shared so agencies can propose against reality. An agency forced to propose blind can only offer generic boilerplate, and you lose the ability to judge their thinking.

The RFP document structure

A good RFP is built to elicit comparable, signal-rich responses. Standard sections:

  1. Overview and context: your business, market, and why you are engaging now.
  2. Objectives and success criteria: the business outcomes above, with how success will be measured.
  3. Scope in and out: what you want the agency to own and, just as important, what you do not.
  4. Requirements and constraints: platform, compliance, reporting, timeline, integration with internal teams.
  5. Proposal instructions: format, length, and the specific questions you want answered, so responses are comparable rather than freeform brochures.
  6. Evaluation criteria: tell them how you will score. Transparency here attracts serious agencies and filters out the ones who will not invest in a real response.

Questions that separate strong from weak answers

This is where selection is won. The objective is questions that are hard to answer well without genuine capability, so a weak agency cannot bluff through them. Across four areas:

  • Strategy. “Given this limited snapshot of our site and market, which three opportunities would you prioritize first, and why?” A strong answer names specific, defensible hypotheses and reasons from the data you gave them. A weak answer recites a generic methodology or asks for a discovery phase before committing to anything.
  • Technical. “Walk us through how you would diagnose a JavaScript site where content is not getting indexed,” or “how you would approach a Core Web Vitals problem.” A strong answer reasons about rendering (the two-wave crawl-then-render model, checking the rendered DOM, server-side rendering options) and names the current thresholds correctly: LCP at or under 2.5 seconds, INP at or under 200 milliseconds (INP replaced FID as a Core Web Vital in March 2024), CLS at or under 0.1, each measured at the 75th percentile of field data. An agency that still talks about FID, recommends AMP for ranking, or points at retired tools is working from a stale playbook.
  • Process. “What do your first 90 days look like?” and “How do you handle a situation where your recommendation conflicts with what the client’s engineering team wants to build?” Strong answers describe a concrete sequence and a real account of navigating organizational friction. Weak answers are vague timelines and conflict-avoidant non-answers.
  • Results. “Tell us about an engagement that failed or underperformed, and what you learned.” This is the single most revealing question, because it is the hardest to fake. A strong agency gives a specific, honest account with a real lesson. A weak one claims never to have failed, which is either dishonest or a sign they have never taken on hard work.

The pattern across all four: strong answers are analytical, specific, and willing to concede uncertainty or failure; weak answers are generic, deflecting, and allergic to admitting a miss.

Weighted scoring across multiple evaluators

Score with a rubric, not a gut feeling, and have three to five people score independently before they compare notes, so one charismatic pitch cannot dominate the room. Weight the criteria by what actually predicts success for your situation. An example template, to be adjusted to your priorities rather than copied as a standard, might weight strategic thinking and technical capability most heavily, with process, relevant results, cultural fit, and price each carrying a smaller share. The weights are yours to set; the discipline is that they are set in advance, applied consistently, and that each evaluator documents the rationale behind a score rather than just a number. Documented rationale is what lets you defend the decision later and what catches the evaluator who scored on vibes.

Reference checks that actually check

Provided references are curated to say good things, so treat them as the floor, not the evidence. Two layers:

  • Scrutinize the provided list. Ask references concrete, hard-to-coach questions: what went wrong and how the agency handled it, whether results were sustained after the engagement, how they responded when an approach was not working, whether they would hire them again for a harder problem. The texture of a real answer differs from a coached endorsement.
  • Add independent research. Find clients not on the list through case studies, LinkedIn, or industry contacts, and look for departed clients. Why someone left an agency is often more instructive than why someone stayed.

Red flags and the contract clauses that matter

Some signals reliably predict a bad engagement. Treat these as disqualifying or at least as demanding a hard explanation:

  • Ranking guarantees (“we’ll get you to position one”). Nobody controls Google’s ranking system, so a guarantee is either ignorance or dishonesty.
  • Vague methodology they will not explain (“proprietary techniques”).
  • Opaque pricing with no line of sight into what you are paying for.
  • The senior team who pitched disappearing after signing, replaced by juniors (team turnover during the pitch is a preview).
  • An inability to name a single failure.
  • Hard-sell pressure and artificial deadlines.
  • A copy-paste proposal that does not engage your specific situation.

And before signing, the clauses that matter more than the monthly fee: termination terms and notice period, intellectual-property ownership of content and deliverables produced, any exclusivity or non-compete restrictions, and liability and indemnification. The differentiator across this whole process is designing the RFP to produce signal (prioritized hypotheses from incomplete data, a real failure) and scoring it with a rubric across several evaluators, so that capability rather than salesmanship decides. A strong process makes a good outcome far more likely; it does not guarantee one, and any agency promising that it does has just shown you a red flag.

Sources