Hiring SEO Talent: Interview Questions and Skill Assessment Methods
On this page
- Define the role against your actual skill gaps first
- Reward diagnostic process over recall in technical questions
- Use a prioritization scenario to test judgment
- Design the practical exercise to be fair and bounded
- Run behavioral questions and references to surface limits
- Frequently Asked Questions
- Are certifications worth anything when screening SEO candidates?
- How do you evaluate a candidate’s claimed results when attribution is ambiguous?
- Should the same interview be used for in-house and agency roles?
- Sources
- Related posts:
SEO hiring is hard for a specific structural reason: the field has no reliable credential, and organic results are influenced by so many factors that a candidate can sincerely claim growth they did not cause. A resume line that reads “drove 30% organic growth” is unfalsifiable in an interview, because you cannot separate the candidate’s contribution from a favorable algorithm update, a brand campaign, a competitor’s stumble, or seasonal demand. The defense is an assessment that probes how a candidate diagnoses and prioritizes rather than what they recall, because process is far more predictive of on-the-job performance than memorized facts, and process is much harder to bluff.
That reframes the whole interview. You are not checking whether someone knows the definition of canonicalization; that is a search away. You are checking whether, handed a real problem with incomplete information, they reason through it in an order that surfaces the likely cause efficiently. Recall is cheap and decays; diagnostic reasoning is the durable skill.
Define the role against your actual skill gaps first
Before sourcing anyone, audit what the existing team can and cannot do. SEO competence splits into roughly four areas that rarely concentrate in one person: technical (crawl, render, indexation, site architecture, performance), content (intent, briefs, editorial quality), links and authority (digital PR, acquisition, risk), and analytics (measurement, forecasting, attribution). A strong technical SEO who cannot evaluate content quality is not a weaker hire than a generalist; they are a different hire, and which one you need depends entirely on where your current gap is.
Writing the role from a genuine gap audit prevents the most common hiring failure, which is hiring another version of who you already have because that profile is familiar and easy to evaluate. It also tells you what to weight in the interview. If the gap is technical, the practical exercise should be a technical diagnosis; if the gap is content strategy, it should be a prioritization-and-brief exercise. The role definition is the rubric’s foundation.
Reward diagnostic process over recall in technical questions
The most predictive single question in an SEO interview is an open diagnostic walk-through, and the canonical version is: “A page you published three weeks ago still isn’t appearing in search. Walk me through how you’d diagnose it.” It cannot be bluffed, because a strong answer has a structure a weak one lacks.
A strong candidate moves through the funnel in a logical order without prompting. They confirm the page is actually published and reachable, then check crawlability (robots.txt disallow, server response, internal links pointing to it). They check for an explicit indexing block (a noindex meta tag or X-Robots-Tag header, which is the single most common silent cause). They use URL Inspection in Search Console to see whether the URL is known, crawled, and indexed, and to view the rendered HTML, since a page that depends on client-side rendering may look complete in the browser but render empty to the crawler.
From there, they consider canonicalization (the page declaring or being clustered to a different canonical), check for a manual action, and only then weigh quality and duplication, recognizing that Google can crawl a page and decline to index it because it judges the content thin or duplicative. A weak candidate says “I’d check if Google crawled it” and stops, with no concept of the rendering step, the selection-versus-crawling distinction, or the order in which causes are most likely.
You are scoring the sequence and the coverage, not any single term. Did they separate “not crawled” from “crawled but not indexed”? Did they mention rendering at all? Did they go from cheap, common causes to expensive, rare ones, or did they jump straight to “the content must be low quality”? The order reveals whether they actually diagnose or merely list things they know.
A second technical probe that separates current practitioners from stale ones touches Core Web Vitals: the three metrics are Largest Contentful Paint (good at 2.5 seconds or less), Interaction to Next Paint (200 milliseconds or less, which replaced First Input Delay as the responsiveness metric in March 2024), and Cumulative Layout Shift (0.1 or less), each assessed at the 75th percentile of real-user field data.
A candidate who still names FID, or who treats lab scores as the thing Google measures, is working from an outdated mental model. You are not testing memorization of the thresholds; you are testing whether they keep current and understand that the assessment is field-based.
Use a prioritization scenario to test judgment
Technical diagnosis tests one muscle; resource judgment tests another. Give the candidate a constrained scenario: limited hours this quarter, a backlog containing a high-volume but very competitive term, a cluster of striking-distance pages sitting on page two, a technical issue affecting a template, and a stakeholder’s pet project. Ask what they do first and why.
A strong answer reasons about expected return against effort and time-to-impact: the striking-distance pages and the template fix likely pay back faster and more reliably than a from-scratch assault on a competitive head term, and the stakeholder request gets weighed on merit rather than reflexively obeyed or dismissed. A weak answer either has no framework (“I’d start with the high-volume keyword because it’s the biggest”) or cannot articulate why one item beats another. You are listening for structured tradeoff reasoning, not the “correct” answer, because the scenario has several defensible orderings and only the reasoning distinguishes them.
Design the practical exercise to be fair and bounded
Interviews sample reasoning; a practical exercise samples work. Keep it tightly time-boxed and give a clear rubric in advance so the candidate knows what success looks like and you evaluate everyone against the same criteria. A good exercise mirrors the role’s actual gap: a short technical audit of a sample site with specific findings and recommended priorities, or a brief plus content outline for a given target query, depending on what you are hiring for.
Two disciplines keep it honest. First, scope it to roughly two hours of effort, and pay candidates for anything beyond that; asking for a full unpaid audit is both unfair and a signal that filters out the strong candidates who have options. Second, score the reasoning and prioritization in the deliverable, not surface polish, because a candidate who flags the three issues that matter and explains why is more valuable than one who lists twenty findings with no sense of which move the needle.
Run behavioral questions and references to surface limits
Use structured behavioral questions in STAR form (situation, task, action, result) so answers are concrete rather than aspirational. The most revealing is the failure question: “Tell me about an SEO initiative that didn’t work and what you learned.” A candidate who says they’ve never had something fail is either inexperienced or not self-aware, and either way that answer is the red flag, because everyone who has done real SEO has watched a confident bet go nowhere. You want a specific failure, an honest account of their part in it, and a concrete lesson applied since.
Structure references to surface limitations rather than collect praise. Ask what the person was best at and where they needed the most support, and ask the referee to describe a situation the candidate found genuinely difficult. Open-ended, limitation-seeking questions get past the reflexive endorsement that confirmation-seeking questions invite, and they tell you where the candidate’s growth edge is so onboarding can target it.
Frequently Asked Questions
Are certifications worth anything when screening SEO candidates?
Treat them as weak positive signals at best and never as a filter. There is no industry-standard credential that reliably predicts capability, and the platform-specific certificates that exist mostly confirm someone completed a course. Diagnostic reasoning and a defensible work sample tell you far more than any certificate on a resume.
How do you evaluate a candidate’s claimed results when attribution is ambiguous?
Don’t try to verify the number; probe the reasoning behind it. Ask how they isolated their contribution from algorithm updates, brand effects, and seasonality, what they would have measured to be confident, and what they would have done if the result had gone the other way. A candidate who acknowledges attribution difficulty is more credible than one who claims clean causation.
Should the same interview be used for in-house and agency roles?
Keep the diagnostic and prioritization core the same, but shift the emphasis. Agency roles lean harder on managing multiple contexts, communicating to non-expert clients, and ramping quickly on unfamiliar sites; in-house roles lean on stakeholder navigation and sustained depth on one property. Adjust the behavioral questions accordingly rather than changing what you test for technically.
Sources
- Google Search Central, Fixing indexing issues / URL Inspection Tool documentation: https://developers.google.com/search/docs/crawling-indexing/url-inspection-tool
- Google Search Central, Block indexing with noindex: https://developers.google.com/search/docs/crawling-indexing/block-indexing
- web.dev, Web Vitals (LCP, INP, CLS thresholds and field measurement): https://web.dev/articles/vitals