Search Central Live Barcelona field trip: what Google told us about crawling, indexing and AI

Jump to section [Open]

    Last week I spent three days in Barcelona at Search Central Live Deep Dive, the most technical event Google runs for SEOs. Each day covered one stage of Search: crawling, then indexing, then serving. I recorded most of the talks on my phone and came home with about 140 photos of slides. 🥴

    I went with one question from our daily work at Bazoom. When a link goes live on a publisher’s site, what has to happen before it counts for anything, in Google or in an AI answer?

    And here’s what I brought back!

    1. AI answers come from the same index as classic search

    Day 2 opened with this slide. Crawled is not indexed: six processing steps sit in between.

    A page reaches the index only if it survives all six steps.

    The part I underlined twice: AI Overviews and AI Mode have no separate pipeline. Google’s documentation says a page must be indexed and eligible for a snippet to appear in AI features. No extra requirements.

    So every indexing problem is also an AI visibility problem.

    2. Snippet controls are now AI controls

    One line from John Mueller’s robots meta session got the whole room typing: nosnippet also stops your content being used as direct input for AI Overviews and AI Mode.

    He also flipped a common habit. Two directives make you more visible, not less: max-snippet:-1 and max-image-preview:large. Use both.

    If robots.txt blocks a page, Google never sees its noindex, and the URL can still appear.

    A new Search Console setting now lets you opt out of AI features. It’s on by default, and opting out means no AI traffic.

    3. Googlebot doesn’t click and doesn’t scroll

    Rebecca Yu from JAKALA gave the most practical 10 minutes of the event. Content that loads on scroll or click doesn’t exist for Google. Her line: “Not visible versus hidden. Present versus absent in the DOM.” A CSS-hidden tab is fine. A tab that fetches its content on click is empty.

    Links follow the same logic. Google extracts a plain <a href>. It doesn’t extract onclick buttons, #/ routes or <span href>. One example from stage: a language selector built as a button left a whole set of alternate-language pages orphaned.

    4. Raw HTML and rendered HTML can tell different stories

    Sören Bendig from Audisto showed real sites where the raw HTML looked perfect and the rendered page was an error message. A travel site listed 312 offers in raw HTML and 200 after rendering.

    If the renderer can’t fetch it, it can’t run it. And robots.txt is per host, so an API subdomain needs its own file.

    On AI, Erin Sparling said AI Overviews and AI Mode handle JavaScript the same as Search. For other AI systems that ground answers on the live web, you probably want server-side rendering. My own note, not Google’s: most third-party AI crawlers don’t run JavaScript at all.

    5. Your canonical tag is a vote, not an order

    “Sometimes.” Yes, someone shipped {{ CANONICAL_TARGET }} to production.

    Google groups similar pages into clusters and picks one URL to represent each. John Mueller was honest about it: rel=canonical is used to cluster, but only sometimes, because many canonicals are bad. Candidates compete on signals weighted by machine learning. Redirects, internal links and sitemaps all vote. Keep them saying the same thing.

    6. Index selection is strict, and blocking first makes it worse

    Gary Illyes shared the number of the week: Google’s systems find 40 billion spammy pages every day. They’re deprioritised before they reach the index. On “Crawled, currently not indexed”, his answer was that most of the time it’s a quality issue, and there’s no need to resubmit.

    A case study from the indexing lightning talks: parameter URLs leaked to Google from JSON inside a script tag.

    No link pointed to them. Blocking them in robots.txt turned them into “Indexed, though blocked by robots.txt”. The fix was boring: remove the source, let Google recrawl, block only as a last resort.

    7. “We like links”, and AI still runs on them

    Links get extracted at the very first processing step, before any rendering.

    Cherry Prommawin’s talk on how Google reads HTML had one slide every link builder should screenshot: “We like links.” Google uses them to discover pages, understand site structure and rank.

    On Day 3, a Google speaker called PageRank a personal favourite while admitting Google leans on it less than it used to. It’s one signal among hundreds now, but still a core system. Link spam got a similar note: less of a focus today, with scaled content abuse the bigger target. My read: the bar moved from link volume to placement quality.

    The AI details worth knowing:

    • Fan-out queries are ordinary queries. AI Mode’s sub-queries hit the same index and get ranked the same way. John Mueller advised against chasing individual fan-out queries, and pages built to game them fall under the spam policies.
    • AI crawlers read content, not link attributes. Mueller said AI crawlers don’t really know what to do with a nofollow link. A brand mention in a strong article is visible to them either way.
    • Brand shows up in the interface. Preferred sources, “Highly Cited” labels and Search profiles all reward sites people already know and trust.

    What this means for link building

    Most talks came back to one question we ask about every placement: does the page get processed well enough to matter?

    If you buy links: a link only passes anything if the page it sits on is crawled, rendered and selected for the index.

    Before you judge a placement by DR, check that the article is indexed, the link is a real <a href> in the raw HTML, the page isn’t noindexed or blocked, and the URL doesn’t hide behind a #.

    And to be honest, we’ve rejected placements on big domains for exactly these reasons.

    Head of SEO @ Bazoom, international speaker, and advocate for smarter link building.