Break Through the AI Noise: Optimize, Be Understood, & Get Cited by AI with Lumar’s GEO Toolkit. Read more.
DeepcrawlはLumarになりました。 詳細はこちら

Building Machine Learning Workflows with Your Crawl Data [Lumar Webinar Replay]

Learn how to move beyond prompt engineering and start building repeatable AI systems that combine machine learning and LLMs.

Lumar SEO Webinar Replay
Watch now:

In our recent webinar, Sam Torres (Sr Manager, Tech SEO, Pipedrive) and Matt Hill (Senior Solutions Engineer, Lumar) outlined how to move beyond prompt engineering and start building repeatable AI systems that combine machine learning and LLMs to solve real SEO problems at scale.

In the session they outlined:

  • Why prompting alone becomes a dead end for repeatable SEO work
  • The three-layer framework combining embeddings, clustering and LLMs
  • How to identify competitor content gaps with far more context than traditional keyword analysis 
  • How semantic clustering can uncover internal linking opportunities that keyword matching misses
  • Practical ways to automate analysis while maintaining quality and consistency

Read on to learn more about the key takeaways read one or alternatively watch the full session above.

The challenge

Your tools give you data. Acting on it at scale is where most teams stall.

Matt Hill identifies that most of the work we are doing in the AI search era is about one-shot prompts. These aren’t connected to our tech stack in a meaningful way. And they aren’t scalable to the point of adding value to the whole system.

This isn’t a new problem, as Hill notes, from onboarding new SaaS systems to rolling out new technologies to our colleagues – repeatable frameworks need to be put into place to get the most out of our data as we navigate through the AI era.

Simply put, a repeatable framework: saves time, maintains quality, and applies across different business needs.

“Because search is at the heart of so much,” Sam Torres says. “There’s so much that search can be applied to outside of just marketing. You can use those insights for product development and all different kinds of things.”

For Torres, machine learning is fundamental to building this framework. They have been subject to academic testing. A lot more of the biases have been defined. But ultimately, machine learning is repeatable.

“You get the same output pretty much every time because it’s not the black box that an LLM is,” she adds. “It also means that these things are documented.”

Using machine learning alongside LLMs is the best of both worlds. They are cost-effective and make for better workflows.

The framework

Embeddings → Clustering → LLMs

For Torres, the first step within her framework is embeddings. This involves taking the content from our pages, giving each piece coordinates, and placing them into a map – producing vectors. 

From there we can manipulate that content using more math and more models.

The next phase is clustering. Using the map to see which pieces of content are close to each other, which are related, and to surface patterns.

Torres’ third phase is the LLM phase – the labelling of this page content. This keeps things organized and easier to work with, turning cluster output into actionable decisions.

“Each step is powerful, but they definitely need each other to be meaningful,” Torres says. “Because if you have similarity scores, that doesn’t always help, right? We can have relevant scores, but you still need to be able to action that. You still need to be able to dig in and see: What does that mean?

Torres’ 4 tips before building workflows

  1. Direct and delegate like you have an intern – a bright intern who’s never seen your stack, your brand, or your stakeholders. Capable but new. The clearer your brief, the better the work.
  2. Think 80/20 for effort – this can mean 20% of the work, delivers 80% of the result. Torres forces us to ask whether we are sacrificing making change on waiting for perfection? Getting 80% of the way there, for 20% of the effort is a win.
  3. Refer to the Is it worth the time? comic – how often you do the task and how much time it takes should determine how long you spend building an automated tool for it. [Note: this is also available as an app at: isitworththetime.com]
  4. Start small, build momentum – if it helps your life a little bit, it’s worth it. It doesn’t have to do it all from day one! Enjoy the wins.

Competitor gap workflow

Hill identifies that when it comes to competitor gaps, we’ve moved away from the Which keywords am I missing? question to: Who owns this topic space, and where are the gaps in breadth and depth?

The first phase of the workflow for Torres, here, focuses on our inputs:

  • Your crawl report (URLs + page content)
  • Competitor URLs (crawled or sampled)
  • GSC data (optional – intent signal)

These inputs are about turning pages into math, with every page becoming a point in semantic space.

“Once we make it math, once we make it numbers,” she says, “it’s really easy to start running analysis and equations on it.”

From there, we can cluster topics and pages. Torres thinks of these as neighbourhoods

You can see where competitors have clusters and you don’t. You can see where your clusters are thin compared to theirs.

We can then move onto labelling, identifying and prioritizing.

LLMs can be used to label the clusters into human-readable terms. They can then flag missing coverage, thin coverage, or where there is over-investment. And they can then rank opportunities by business impact. Not volume!

The result? A prioritized gap report, grounded in structure.

Internal linking workflow

For Hill, internal linking is still seen to be a manual thing – leading to inconsistent results.

Torres identifies that linking sits at the intersection of 4 functions:

  1. CMS
  2. Content strategy
  3. Crawl data
  4. Engineering

Her workflow has the same layers as the competitor gap method, but with a different direction of analysis. 

Whereas competitor gap is looking outward at content from competing brands and the external topic landscape, the internal linking workflow is looking inward at our own content graph, our own neighbourhoods, and the linking opportunities that are already there.

Again, we start with our inputs. In this case, the required input would be our crawl export. Another optional input would be an existing link map, which lets you score what’s already covered.

We can then put our content and link graphs into vectors – with the clusters surfacing semantic relationships that anchor-text matching would miss. Crucially, we can see pages that belong together even when they share no keywords.

“So what’s cool about that is the system can analyze all of that and then it can also make recommendations,” Torres says. “So where should links be built?”

From there it is possible to build links with the highest confidence scores, improving your link-building efficiency, and allowing for better use of your time in other areas.

Implementation

In summary, Hill reminds us that while AI is everywhere, scale isn’t. 

Both of Torres workflows can really help build the difference.

Torres’ closing remark comes back to how powerful machine learning models are in terms of scaling up our work, as well as being cost-effective.

“Machine learning models for translation outperform LLMs pretty much every time,” she says. “So if you’re thinking about those kinds of things… I would encourage you to think about your workflows. Where can you integrate a machine learning model?”


Don’t miss the next Lumar webinar!

Sign up for our newsletter below to be alerted about upcoming webinars, or give us a follow on LinkedIn to stay up-to-date with all the latest news in SEO, GEO, and digital optimization.

Want even more SEO insights on-demand? Browse Lumar’s full library of SEO and website optimization webinars here.

Newsletter

Get the best digital marketing & SEO insights, straight to your inbox