20% Offfor startup website packages

Digital Marketing

Technical SEO Checklist (2026): 25 Checks You Can Do Yourself

By Webhorse Studio Editorial Team

25 technical SEO checks you can do yourself with free tools, from indexing and sitemaps to Core Web Vitals, schema and AI crawler access.

Cover image for Technical SEO Checklist (2026): 25 Checks You Can Do Yourself

A technical SEO checklist is the list of site-level checks that decide whether Google can find, crawl, render and index your pages before content quality even comes into play. The 25 checks below cover indexing, sitemaps, robots.txt, canonicals, redirects, Core Web Vitals, mobile, HTTPS, structured data, JavaScript and AI crawler access. Every one of them can be done with free tools, mainly Google Search Console and PageSpeed Insights.

We run this list on new client sites before launch and on older sites when traffic drops. It is written for business owners and small marketing teams, so each check says what to look at, where to look and what "fixed" means. A priority table at the end tells you what to tackle first.

What technical SEO covers (and what it doesn't)

Technical SEO deals with how your site is built and served: URLs, status codes, crawl rules, speed, rendering and markup. On-page SEO deals with what a single page says, such as its title, headings, copy and internal links. Off-page SEO is mostly links and mentions from other sites.

The order matters. A page that Google cannot crawl or index will not rank however good the writing is, so technical problems are worth fixing first. Once the basics are clean, most of the ranking gains come from content and links. If you want the bigger picture, our older guide on getting your website on Google's first page covers the content side.

Free tools you need for this technical SEO checklist

ToolCostWhat you use it for
Google Search ConsoleFreePage indexing report, URL Inspection, sitemaps, robots.txt report, Core Web Vitals report, HTTPS report
PageSpeed InsightsFreeReal-user Core Web Vitals (field data) and a Lighthouse lab test per URL
Rich Results Test / Schema Markup ValidatorFreeChecking structured data for errors
Chrome DevToolsFree, built into ChromeViewing status codes, redirects, mixed content and rendered HTML
Bing Webmaster ToolsFreeA second opinion on crawl and index issues, and coverage for Bing
A desktop crawler (free tier)Free up to a page limitFinding broken links, redirect chains, missing titles and orphan pages across the whole site

Set up Search Console first. Most checks below start there, and some reports need a few days of data after you verify the site.

Crawling and indexing (checks 1 to 7)

1. Confirm your important pages are indexed

Open Search Console's Page indexing report. It splits your URLs into indexed and not indexed, with a reason for each excluded group. Google lists reasons such as "Crawled - currently not indexed", "Discovered - currently not indexed", "Duplicate without user-selected canonical" and "Soft 404" (Google: Page indexing report).

Not every excluded URL is a problem. Redirects, filtered URLs and pages you deliberately set to noindex should be excluded. What you are looking for is a service page, product or blog post that should rank but sits in the "not indexed" list.

2. Inspect your top pages one by one

Paste your home page and your five to ten most valuable URLs into the URL Inspection tool. Check that each is indexed, that the "Google-selected canonical" matches the URL you expect, and that the page was crawled by the smartphone crawler. If a key page is not indexed, fix the cause first and then click "Request indexing".

3. Check robots.txt is not blocking anything important

Visit yourdomain.com/robots.txt. Make sure no Disallow rule covers pages you want in search, or the CSS and JavaScript files those pages need to render.

Robots.txt controls crawling, not indexing. Google says it is "not a mechanism for keeping a web page out of Google" and that a blocked URL can still be indexed if other pages link to it (Google: robots.txt introduction). To keep a page out of search, use a noindex tag or a password.

4. Make sure noindex tags are only where you want them

A leftover <meta name="robots" content="noindex"> from a staging site is a common reason a newly launched website gets no search traffic. Look for "URL marked 'noindex'" in the Page indexing report and check each URL listed. On WordPress, also check that "Discourage search engines from indexing this site" under Settings, Reading is unticked.

Do not combine a robots.txt block with a noindex tag on the same page. If Google cannot crawl the page, it never sees the noindex.

5. Submit a clean XML sitemap

Your sitemap should list only the canonical, indexable URLs you want in search, written as full absolute URLs. A single sitemap file is limited to 50,000 URLs or 50MB uncompressed, and Google ignores the <priority> and <changefreq> values. It does use <lastmod>, but only if the dates are consistently accurate (Google: build a sitemap).

Submit the sitemap under Sitemaps in Search Console and check the status says "Success". Remove redirected, noindexed and 404 URLs from it.

6. Fix server errors and soft 404s

Status codes tell Google how to treat a URL. Google treats 404 and 410 the same way and drops the URL from the index. Repeated 5xx server errors make Google slow down crawling, and indexed URLs are eventually dropped if the errors continue. A "soft 404" is a page that returns 200 OK but looks like an error or empty page (Google: HTTP status codes and network errors).

Check the Page indexing report for "Server error (5xx)" and "Soft 404". Empty category pages, "no results" search pages and out-of-stock product pages are frequent soft 404s.

7. Keep important pages close to the home page

Pages that are buried deep, or that no other page links to (orphan pages), get crawled less often and pass less authority. Run a crawl and sort by click depth. Your main service pages and best-selling categories should be reachable from the main menu or the home page, and every page in your sitemap should have at least one internal link pointing to it.

URLs, duplicates and redirects (checks 8 to 12)

8. Pick one version of your domain

Your site can usually be reached four ways: http and https, with and without www. Choose one and permanently redirect the other three to it. Type each version into your browser and confirm they all land on the same address in a single hop.

9. Use canonical tags on duplicate and filtered pages

Ecommerce filters, tracking parameters and printer versions create many URLs with the same content. Add a rel="canonical" tag that points to the main version. Google ranks canonicalization signals by strength: a redirect is a strong signal, rel="canonical" is a strong signal, and sitemap inclusion is a weak one. Google also advises against using robots.txt or noindex to choose a canonical (Google: consolidate duplicate URLs).

In URL Inspection, if "Google-selected canonical" differs from "User-declared canonical", Google disagrees with your choice. That usually means the two pages are too similar or your internal links point to the other version.

10. Use permanent redirects for permanent moves

When a page moves for good, use a 301 (or 308) redirect. Google treats a 301 as a strong signal that the target should be indexed and a 302 as a weak one. Check which type your CMS or redirect plugin uses.

11. Remove redirect chains and loops

A chain is A redirects to B, which redirects to C. Each hop slows users down and wastes crawling. Point A straight to C, and update internal links so they go to the final URL directly. Your crawler's redirect report lists every chain. Site redesigns and http-to-https moves leave a lot of them behind.

12. Fix broken internal links

Internal links to 404 pages waste crawl activity and send visitors to dead ends. Run a crawl, export the list of 4xx links and either update each link or redirect the dead URL to the closest live page. Keep URLs short, lowercase and hyphenated when you create new ones, and avoid changing them once they rank.

Speed and Core Web Vitals (checks 13 to 16)

13. Check real-user Core Web Vitals

Core Web Vitals are three metrics. Largest Contentful Paint (LCP) measures loading and should be 2.5 seconds or less. Interaction to Next Paint (INP) measures responsiveness and should be 200 milliseconds or less. Cumulative Layout Shift (CLS) measures visual stability and should be 0.1 or less. Each is judged at the 75th percentile of visits, separately for mobile and desktop (web.dev: Web Vitals).

INP replaced First Input Delay (FID) on 12 March 2024 (web.dev). Older checklists that still mention FID are out of date.

Use the Core Web Vitals report in Search Console for groups of pages, then PageSpeed Insights for individual URLs.

14. Read field data and lab data separately

PageSpeed Insights shows two kinds of results. Field data comes from the Chrome User Experience Report and covers real visits over the previous 28 days. Lab data comes from a simulated Lighthouse test. New or low-traffic pages may not have field data, in which case PSI falls back to data for the whole site or shows none (Google: About PageSpeed Insights).

Field data is what Google's assessment is based on. Use the lab score to find causes and test fixes, and don't chase a perfect 100.

15. Fix the usual LCP and CLS causes

Most slow LCP on small business sites comes from a large hero image or slider, slow hosting, or render-blocking scripts. Most CLS comes from images, ads and embeds without set width and height, or from fonts and banners that load late and push content down.

  • Compress hero images and serve WebP or AVIF at the size actually displayed.
  • Do not lazy-load the main image at the top of the page.
  • Set width and height attributes on images, videos and iframes.
  • Remove plugins, sliders and chat widgets you don't use.
  • Use caching and a CDN if your host offers one.

16. Cut heavy JavaScript to improve INP

Poor INP usually means the browser is busy running scripts when someone taps a button or opens a menu. Tag managers loaded with old tags, multiple analytics tools, chat widgets and page builders that ship large scripts are the usual causes. Remove what you don't need and delay non-essential scripts until after the page has loaded.

Keep this in proportion. Google says "Core Web Vitals are used by our ranking systems" but also that a good score doesn't guarantee top rankings, and it will still show the most relevant page even if its page experience is weaker (Google: page experience).

Mobile and security (checks 17 to 19)

17. Make the mobile version complete

Google indexes and ranks using the mobile version of your content, crawled with its smartphone agent. Its guidance is to keep the same primary content, the same structured data and the same robots meta tags on mobile as on desktop, and not to lazy-load primary content that only appears after a user clicks or swipes (Google: mobile-first indexing best practices).

Open your key pages on a phone. Check that text is readable without zooming, buttons are easy to tap, and nothing important from the desktop page is missing.

18. Remove intrusive pop-ups

Google's page experience self-check asks whether your pages avoid intrusive interstitials. A full-screen newsletter or app-install pop-up that covers the content on mobile is the classic example. Use a small banner instead, or show the pop-up only after the visitor has scrolled or spent time on the page.

19. Serve every page over HTTPS with no mixed content

All pages should load over HTTPS, with http URLs redirecting permanently. Then check for mixed content: images, scripts or fonts still loaded over http on an https page. Chrome DevTools' Console shows mixed-content warnings, and Search Console has an HTTPS report. Our website security checklist covers certificates, updates and backups in more detail.

Structured data, JavaScript and international (checks 20 to 22)

20. Add structured data Google still uses

Structured data (usually JSON-LD) helps Google understand what a page is about and can make it eligible for rich results. Google's current search gallery lists types including Organization, Local business, Breadcrumb, Article, Product, Review snippet, Event and Video (Google: structured data search gallery). For most business sites, Organization or LocalBusiness on the home or contact page plus Breadcrumb sitewide is a good start.

Note for 2026: Google retired FAQ rich results. Its documentation states the feature stopped appearing in Google Search from 7 May 2026 (Google: FAQPage). FAQ sections are still useful for readers, but don't expect FAQ markup to add expandable answers under your result. Validate any markup with the Rich Results Test, and only mark up content that is visible on the page.

21. Check that JavaScript content is crawlable

Google processes JavaScript pages in three phases: crawling, rendering and indexing. It recommends server-side rendering or pre-rendering because it is faster for users and crawlers, and "not all bots can run JavaScript". Links need to be normal <a href> elements for Google to follow them (Google: JavaScript SEO basics).

To test, use URL Inspection, click "Test live URL" and view the rendered HTML. If your main text, prices or menu links are missing there, Google is not seeing them either. React, Angular and Vue sites built as single-page apps are the ones to check most carefully.

22. Set up hreflang if you have more than one language

If your site has, for example, English and Hindi or Marathi versions, use hreflang to tell Google which version to show to whom. Google supports ISO 639-1 language codes with optional ISO 3166-1 region codes, such as en-IN or hi-IN. Each version must list itself and every other version, or Google may ignore the tags, and x-default covers visitors who match none of them (Google: localized versions). Single-language sites can skip this check.

AI crawler access (checks 23 to 25)

23. Decide which AI crawlers to allow

AI search tools use their own crawlers, controlled through robots.txt. These are the ones most site owners need to know:

User agent / tokenWhat it controlsEffect of blocking it
Google-ExtendedWhether Google can use your content to train future Gemini models and for grounding in Gemini apps and Vertex AIGoogle says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal" (source)
OAI-SearchBotSurfacing your site in ChatGPT's search featuresYour pages won't appear in ChatGPT search answers, except as navigational links (source)
GPTBotContent that may be used to train OpenAI's foundation modelsSignals your content shouldn't be used for training; does not control ChatGPT search
ChatGPT-UserVisits a page when a ChatGPT user asks about itOpenAI says it is not used to decide whether content appears in search

For most businesses that want enquiries, allowing OAI-SearchBot makes sense. Whether to allow training crawlers such as GPTBot and Google-Extended is a business choice, and blocking them does not remove you from Google Search. OpenAI notes robots.txt changes can take about 24 hours to apply to its search.

24. Check your firewall or CDN isn't blocking good bots

Some hosting firewalls and CDN "bot protection" settings block crawlers regardless of robots.txt. If Search Console shows unexplained crawl errors, or you have allowed a crawler in robots.txt but it still can't reach you, check your CDN or security plugin's bot settings and logs. OpenAI, for example, asks site owners to allow requests from its published IP ranges as well as allowing OAI-SearchBot in robots.txt.

25. Put key facts in plain HTML

AI tools and search engines read the HTML they receive. Your services, locations, prices (if you publish them), contact details and opening hours should be in real text on the page, not only inside images, PDFs or scripts. Clear headings and short, direct answers near the top of a page make the content easier for both people and machines to quote. This overlaps with checks 20 and 21: get rendering and markup right and you've covered most of it.

Priority order: what to fix first

You don't need to do all 25 at once. This is the order we use, based on how much each problem can cost you.

PriorityChecksWhyHow often to recheck
Fix today3 robots.txt, 4 noindex, 6 server errors, 8 domain version, 19 HTTPSAny of these can remove pages, or the whole site, from searchAfter every launch, redesign or migration
Fix this week1 indexing, 2 URL Inspection, 5 sitemap, 9 canonicals, 10 redirects, 21 JavaScriptDecide which pages Google indexes and which URL it showsMonthly
Fix this month7 click depth, 11 redirect chains, 12 broken links, 13 to 16 Core Web Vitals, 17 mobile, 18 pop-upsAffect crawl efficiency, user experience and close ranking contestsMonthly, and after adding plugins or scripts
Then20 structured data, 22 hreflang, 23 to 25 AI accessImprove how your pages are understood and shown, once the basics are cleanQuarterly

Quick technical SEO audit checklist (copy this)

  • Search Console and Bing Webmaster Tools verified
  • Key pages show as indexed in URL Inspection
  • Google-selected canonical matches yours on key pages
  • robots.txt blocks nothing you want in search
  • No stray noindex tags; WordPress "discourage search engines" unticked
  • XML sitemap submitted, status "Success", only canonical 200 URLs
  • No 5xx errors or soft 404s on important pages
  • One domain version; other three redirect in one 301 hop
  • No redirect chains; no internal links to 404s
  • LCP 2.5s or less, INP 200ms or less, CLS 0.1 or less on mobile field data
  • Mobile pages carry the same content, structured data and meta robots as desktop
  • No full-screen pop-ups on mobile
  • HTTPS everywhere with no mixed content
  • Organization or LocalBusiness and Breadcrumb markup validated
  • Rendered HTML contains your main text and links
  • hreflang set up correctly (multilingual sites only)
  • AI crawler rules in robots.txt match your business decision

When to get help

Most of this list is a few hours of checking for a small site. It gets harder when the fixes need code changes: JavaScript rendering, Core Web Vitals on a heavy theme, large-scale canonical problems on an ecommerce catalogue, or cleaning up after a migration. If you are weighing up an agency, our guide to SEO cost in India explains what audits and monthly retainers usually include.

Webhorse Studio fixes these issues as part of our SEO and digital marketing services, and we build new sites with them handled from day one.

Frequently asked questions

What is included in a technical SEO checklist?

A technical SEO checklist covers how search engines access and process your site: indexing, robots.txt, XML sitemaps, canonical tags, redirects and status codes, site speed and Core Web Vitals, mobile usability, HTTPS, structured data, JavaScript rendering, hreflang for multilingual sites and, in 2026, AI crawler access.

Can I do a technical SEO audit myself with free tools?

Yes, for most small and medium sites. Google Search Console, PageSpeed Insights, the Rich Results Test, Chrome DevTools and the free tier of a desktop crawler cover all 25 checks above. Paid tools save time on large sites but find the same core issues.

How often should I run a technical SEO audit?

Do a full audit before and after any redesign, migration or platform change. Otherwise, check the Page indexing and Core Web Vitals reports monthly and run the full list once a quarter.

What is the difference between technical SEO and on-page SEO?

Technical SEO is about the site's infrastructure: whether pages can be crawled, rendered, indexed and loaded quickly. On-page SEO is about each page's content: titles, headings, copy, images and internal links. Technical problems usually need fixing first, because a page Google can't index can't rank.

Do Core Web Vitals affect Google rankings?

Google says Core Web Vitals are used by its ranking systems, but a good score does not guarantee top positions. Relevance comes first. Speed and stability matter most when several pages are similarly useful, and a slow, jumpy page is harder for visitors to use.

Should I block AI crawlers in robots.txt?

It depends on what you want. Blocking OAI-SearchBot keeps your pages out of ChatGPT search answers, which most businesses don't want. Blocking training crawlers such as GPTBot or Google-Extended stops that use of your content, and Google says Google-Extended does not affect inclusion or ranking in Google Search.

Want a second pair of eyes on your site? Contact Webhorse Studio and tell us your URL and what's going wrong. We'll tell you which of these checks to start with.

Our Newsletter

Get Every Single Updates Newsletter Subscribe

Get Every 7 Days Updates or Monthly updates

Newsletter illustration