Data Matters

Published by

on

Digital data visualization with graphs, charts, globe, and circuitry representing data analytics

When we think of the world of AI today, with seemingly any information instantly at our fingertips, it’s understandable that the excitement is palpable. But I’m reminded of the early days of the internet & some of the bad data that affected it.

The Internet was an amazing opportunity to publish and access data in an instant from anywhere in the world. Just like the desktop publishing revolution that had preceded it, the Internet made it easier for people to share their ideas. We didn’t need a large publisher to decide to publish our book, instead we could post a website and some pages about whatever came to mind. Coming from the world of physical books and your local library, this was ground breaking.

But with so much information available, discovering and locating it became a challenge. Search engines aimed to become the arbiters of which content to surface at the top of the results. Add opportunities for advertising & people tried to game the system by loading web pages with high-demand keywords, as they could earn money from that first click.

The issue is that these pages would appear at the top of search results because they had the most matches of the popular search terms in their content. But it didn’t mean that the content was valuable or even accurate. Google took a different approach with PageRank, in which the number of links pointing to a page provided a signal of its value. So, loading a site with a bunch of the search phrases wouldn’t work anymore.

But, like anytime you measure something, it becomes what you optimize for, sites learned this and started creating link farms. One site would link to others (for a fee) and others to it, to try to game the PageRank system.

Links from sites that themselves had high pageranks carried even more weight. This was an effort to ensure that only high quality sites would be at the top while the link farms did affect this a bit, Google really needed to watch for bad content & effectively discount its weight to avoid bogus, unuseful links from appearing at the top of search results.

While I’ve focused a lot on the mechanics, it’s important to remember the purpose of the search engines was to help humans find content that was relevant and useful for their query. Link farms & key phrase stuffing don’t remove the human intention for using a search engine, if anything that aim to subvert it.

Fast forward to today’s GenAI and we effectively have a new search engine. Unlike search engines of old which would list links to the most relevant results (using things like PageRank to sort them), GenAI returns a document that is a synthesis of all the content the model has absorbed that aligns with the query. Although the returned result is different, it is still susceptible to the same manipulation by bad input data.

With even more AI slop being generated and then published on the web, we hit a circular data issue. The AI is generating output based on its own output (or that of another AI). If we have no editorial or even attempts to tag input data as reputable or questionable, all data is considered as valuable as any other. That naive approach is simple and scalable, but it doesn’t solve what humans are asking for in their prompt.

We focus so much on the technical achievements of modern LLMs, how big the context window is, or how much data the model has been trained on, we forget what really matters about any knowledge system…. the data. There is a reason data scientists have to have some domain expertise, they need to understand what the data actually means & how to interpret it.

In a closed system, using only internal data, this can be manageable & even lead to useful and insightful results when using AI. But when you are handling the wild west, that is the broader internet with so much data growth over time, it is impractical for commercial LLMs to have this human editorial & data evaluation.

While there is an adage that more data is always better, that only holds true when the data is as reliable as the stuff you already have. Lots of bad data, doesn’t give you a better prediction. Lots of bad data, doesn’t yield insights humans couldn’t have found. Lots of bad data doesn’t improve the human experience of using these tools. Lots of bad data just makes everything worse.

In the end, data matters because each piece of data is an ingredient in anything you are creating. The algorithms used are like the techniques a chef uses in preparing a dish. But if you have rotten ingredients, it doesn’t matter how good the techniques are…. you’ll have an inedible, or even poisonous dish. The same goes for software and tools. Bad data will create bad software output. Data matters!

Discover more from Niels Meersschaert

Subscribe now to keep reading and get access to the full archive.

Continue reading