AI Datacenters Are Eating the World

AI datacenters are categorically different from traditional “hyperscale” cloud providers. The older datacenters were optimized for networking and storage – think streaming video and commerce websites.

AI datacenters are optimized for computing, and density is king. The goal is to pack as many parallel processing chips – Nvidia GPUs and Google TPUs – into a rack, and as many racks into the building, as possible.

This means that power consumption and cooling requirements are through the roof. One rack in a typical AWS datacenter might draw 20 kilowatts, while the latest Nvidia rack draws 132 kW. Pack the building full of those, and…

Project/Company Target Capacity Status & Timeline Key Details
xAI Colossus Expansion 1.2 GW Expansion underway; 150 MW substation completing Q4 2025, full 1.2 GW by late 2025–early 2026 Building on the existing 250–300 MW Colossus cluster (built in 122 days). Involves on-site natural gas plant and grid upgrades; faces EPA scrutiny but leverages pre-existing factory infrastructure for speed.
OpenAI Texas Campus 1 GW (phased from 300 MW) Phase 1 (300 MW) operational; Phase 2 construction started Jan 2025, full GW by mid-2026 Houses hundreds of thousands of GPUs; includes 210 substations and on-site electrical upgrades. Already straining ERCOT grid—could equal ~10% of regional peak during heat waves.
Meta 1–5 GW (supercluster) Ground broken 2024; first 1 GW online 2026, scaling to 5 GW by 2030 Zuckerberg’s “gigawatt-plus” initiative; Meta’s largest yet, with $64–72B spend in 2025 alone. Focuses on liquid cooling for high-density AI racks; part of broader multi-GW campus plans.
Microsoft OpenAI “Stargate” 1–5 GW Planning advanced; construction to start late 2025, launch 2028 $500B joint venture with Oracle/Nvidia; aims for massive AI sovereignty. Includes SMR nuclear pilots; power sourcing via PPAs and on-site generation to bypass grid delays.

Microsoft has two 300 MW datacenters. This is comparable to peak load for the city of Tacoma, during their summer AC season. Within a few years, all the leading AI vendors will have datacenters above 1GW. That’s why Microsoft just made a deal to restart the infamous Three-Mile Island nuclear power plant.

A cynic might observe that, while the TMI facility was deemed unsafe to power homes and businesses in Pennsylvania, regulators were willing to reconsider once Microsoft came knocking. Likewise, in Europe, nuclear-powered France is winning the datacenter race over green Germany.

After years of woodburning and windmills, the voracious demands of AI are forcing the world to take another look at nuclear power.

The Art Project

Can AI be used to match and classify images? Of course! They do this all the time, looking at everything from paint chips to x-rays. In today’s post, I use an established model called ResNet-50 to match and classify post-impressionist artists. For example, Braque and Picasso have a 70% similarity score.

The “cosine similarity” between Braque and Picasso is 0.70.

ResNet-50 is a convolutional neural network (CNN) introduced in 2015. Normally, we would use it as the base for image interpretation, and then add layers to learn the specific application. In this case, we are only interested in the coding system it uses, called an “embedding.”

ResNet-50 encodes each image as a list of 2,048 numbers, known as a “vector” in machine learning. This vector is not simply a way to store the image – the JPEG file already does that – but to encode whatever features the model deems useful.

For this demonstration, I collected examples from fourteen artists. To avoid complications over the choice of subject, I used self-portraits by each artist.

Experiments with CNNs show that they recognize shapes, colors, styles, and textures – everything you would expect from “machine vision.” Our model is not going to know anything about the painters, though – not who cut off an ear, or who moved to Tahiti. It’s just the pixels.

With the fourteen paintings vectorized, we can do things like compute similarity scores. For instance, Braque, Chagall, and Picasso seem to hang together. I also ran a hierarchical clustering analysis.

It’s hard to imagine what the clustering algorithm “sees” in high-dimensional space so, wherever possible, I try to reduce down to three dimensions – using principal component analysis (PCA) or UMAP. In this case, because of the small sample, a three-D chart captures 40% of the variance.

The human eye naturally finds clusters – there are Picasso, Braque, and Chagall down at the bottom, and here is Kandinsky off by himself. Also note that Cezanne, Gauguin, and Schiele are spread out along the Y axis, but together on the X axis.

Unfortunately, these axes are completely arbitrary. ResNet-50 can’t tell us if Z is the “axis of cubism,” or whatever. That’s the knock against neural net reasoning being a “black box.” We can see, though, that the PCA plot roughly agrees with the cluster analysis.

So, that was about two hundred lines of code as a proof of concept, plus some fun charts. If you were really doing this for your MFA, you would want to use many more paintings, and stash them in a vector database. For more on vector databases, see Literary Analysis with RAG.

Today’s featured image nods to a common gaffe in generative AI. Yes, Marc Chagall really did paint a “Self-Portrait with Seven Fingers.”

Project Avatar

Everything in this video is AI generated. My voice and image have been cloned. Even the script was generated, by a Google product called Notebook LM. This post is mostly about Notebook LM, and I’ll also survey some other Gen AI tools.

Notebook LM is basically RAG in a box. If you don’t know what that is, you can read my earlier posts on the topic – or you can watch the video. I thought it would be clever to feed RAG articles to a RAG system, and have it generate a dialogue.

That’s right, Notebook LM will ingest raw source material, and then generate a podcast-style dialogue. The system is meant as a study aid, and you can imagine how powerful that is. Other outputs include a study guide, FAQ page, and timeline. Here is a sample entry from the War and Peace timeline:

October 1805: News arrives of Mack’s defeat. The Pavlograd hussars, including Rostov and Denisov, are stationed near Braunau. Rostov experiences his first taste of battle. He witnesses the horrors of war and feels disillusioned. Prince Andrew serves as an adjutant for Kutuzov.

One challenge with RAG has always been preparing the source materials. This earlier post described the work of parsing and vectorizing several text files. In real world applications, clean source material is hard to find. Notebook LM swallows PDF files with ease.

I was curious about the health concerns around seed oils, so I rounded up some papers from sources like the Journal of Nutrition and Metabolism, and just dumped them into Notebook LM. It prepared a handy summary of each one, plus the outputs listed above. I listened to the dialogue and, of course, you can chat with it, too.

  • Source Summaries
  • FAQ Page
  • Study Guide
  • Table of Contents
  • Timeline
  • Briefing Book
  • Chat Window
  • Dialogue

This is a practical, down-to-earth application of LLM technology. One person I found on Reddit is using Notebook LM to prepare for the CISSP exam. He’s doing what I did with seed oils, hoovering up all the InfoSec papers.

From Podcast to Video

Since the Notebook LM dialogue is audio only, I thought it would be fun to make a video and cast my own avatar for the male voice. That’s not even a real photograph of me. First, I trained a photo avatar on HeyGen, and then requested “Mark wearing dress shirt in library.”

Synthesia is similar to HeyGen, but it’s optimized for training videos. It uses a slideshow format. People like it because, if this is your application, all the tools are in one place. I found HeyGen to be more flexible for things like photo avatars and voice substitution.

Other tools I looked at were Deepbrain, now AI Studio, Wondershare, D-ID, and Creatify. Creatify is optimized for making product advertisements on social media. It can write its own script, based on reading the product’s website.

For my voice, I made an “instant voice clone” on Eleven Labs. I didn’t have the patience to make a “professional” one. The instant clone is good enough and, frankly, a little creepy.

I selected a canned avatar named Georgia to be my partner. Initially, I used the script from Notebook LM, and ran the HeyGen animation in text-to-speech mode. Georgia is native there, and HeyGen was able to use my voice via API from Eleven Labs. HeyGen also supports integration with LMNT, Play.ht, and Cartesia.

This is, by far, the easiest way to do it. When it was time to combine the two videos, I was able to use transcript-based editing in CapCut. Unfortunately, the result was a little bit robotic. The charm of Notebook LM’s dialogue is that it really sounds natural.

While one is speaking, the other sits patiently and makes facial expressions as if listening.

Working with the WAV file was more challenging. I used Audacity to split the male and female roles – not easy when they interject “uh-huh” over each other’s lines, but that’s the desired effect.

I left the female voice as-is, ran the male audio track through Eleven Labs to pick up my voice, and then went back to HeyGen – this time, uploading prefab audio for me and Georgia (separately) instead of scripts.

The result from HeyGen was two videos, one for each avatar. While one is speaking, the other sits patiently and makes facial expressions as if listening. The timing works because the split tracks from Audacity are in sync. The last thing to do was combine these, split-screen, in CapCut.

Gen AI and Social Media

My work with AI has always been machine learning for quantitative applications – Python, Scikit, and applied statistics – so it was fun to learn about the crazy things people are doing with Gen AI.

For instance, there is an AI generated influencer on Instagram. An Italian modeling agency created her, so the story goes, because they were tired of working with real prima donnas. 

There is now a cottage industry of avatars on social media, using tools like Creatify to monetize attribution. I thought for a moment about my custom GPT, Powerful Thinking, and its avatar, Bruno. But I couldn’t think of anything for him to sell. Next week, I’ll be back to my regular coding projects.

Literary Analysis with RAG

The situation was dire. Napoleon’s army was far from its supply lines, with the harsh Russian winter closing in. Their only hope of shelter lay in the capital city of Moscow, but, arriving there in September 1812, they found the city in flames. Of the half-million men who had set out with Napoleon, only ten thousand would survive the retreat.

If you ask ChatGPT about this, it will give the conventional answer, plus some convincing – and incorrect – ideas about the book. 

The conventional version of this history says that the Russians purposely torched their own city. Writing in War and Peace, however, Count Leo Tolstoy won’t admit such a desperate tactic. He presents a strong case that Moscow, “a town built of wood, where scarcely a day passes without conflagration” would naturally burn once its people – and the fire department – evacuated.

If you ask ChatGPT about all this, it will give the conventional answer. If you ask, “according to Tolstoy,” it will still give the conventional answer, plus some convincing – and incorrect – ideas about the book. That’s because ChatGPT has never read the book!

It’s adorable, because it bluffs about the reading, exactly like a slacking college student. What was the professor thinking, assigning a thousand-page novel?

Now, thanks to Retrieval Augmented Generation (RAG), you can help ChatGPT answer such questions by priming it with selected passages from the novel. So, I did. I wanted to demonstrate that RAG would support more-advanced text analyses: 

  1. Answer questions using passages from a single novel
  2. Answer questions using passages from two novels and compare them
  3. Compare passages from two novels based on an unseen question
  4. Compare passages from two novels based on a similarity search

This week, we’ll cover the basics using War and Peace, and then I’ll share the two-novel results next week.

War and Peace and RAG

The basic idea behind RAG is simply to query a text database for the priming material, before handing the problem over to a Large Language Model (LLM) like ChatGPT. To be precise, I am using the OpenAI API to work with the GPT 3.5 model.

The only AI involved in the retrieval step is that we use an “embedding model” to convert the text and the query string into vectors. Apart from that, it’s a text search. You could, conceivably, use old-school text search techniques to do the job. I haven’t tried that, yet. What I tried were these three embedding models:

  • text-embedding-3-large
  • text-embedding-ada-002
  • text-embedding-3-small

The embedding models convert chunks of text into vectors of varying length, hence the “large” and “small” size designations. If you’re an AI person reading this, you already know about mapping words to vectors. Mapping paragraphs to vectors follows the same principle. Both are due to Mikolov, et al. See Distributed Representations of Sentences and Documents.

If you’re a language person, well, you won’t be surprised to learn that the English lexicon can be arranged into a spatial array so that “cat” and “dog” end up together. And, if a three-D word space is good, a 300-D word space is better!

I downloaded some novels from the Gutenberg project, did some basic text parsing on them, and then converted each into its own vector database. I used the Chroma database natively and with the LangChain library. Other popular vector databases include Milvus and Pinecone.

I parsed the novels into chapters first, so that Chroma would pick up the chapter headings as metadata. If you use Project Gutenberg, be sure to stop parsing at THE END because there are about 600 lines of legal stuff after that. 

Based on the context provided, it is implied that the Russians did not intentionally set fire to Moscow. The fire in Moscow was attributed to a combination of factors such as the presence of foreign troops, abandoned wooden buildings, soldiers’ carelessness with pipes and campfires, and the general chaos and looting that ensued. The burning of Moscow was seen as a result of the circumstances rather than a deliberate act of arson by the Russians.

The main thing to tune, here, is the size of text chunks in the database. You’ll be passing text chunks to the LLM via an API, and there’s a size limit. GPT 3.5 has a window of roughly 3,000 words.

The typical RAG use case is simply to find one good chunk with the answer. For the Moscow fire question, that is chunk #1440, which happens to be only 0.12 “cosine distance” away from the query text:

{‘source’: ‘data\\tolstoy\\Chapter_255.txt’, ‘start_index’: 10994}

But what if you’re asking “who is the best Russian general?” or “how many siblings does Natasha have?” Clues will be scattered throughout the text, and you will need more, smaller chunks. 

Prince Andrew doesn’t marry Natasha immediately because he decides to conform to his father’s wish to propose and postpone the wedding for a year. This decision is made in the hopes that either his own or Natasha’s feelings may change within that time period. Despite Natasha’s distress and desire to marry sooner, Prince Andrew feels bound by his father’s wish and the decision to delay the wedding.

I would say that smaller is better because the retriever can always fetch multiple chunks from the same neighborhood – as long as they’re at least big enough to be picked up by the vector search. After some experimentation, I settled on a chunk size of 1,200 characters, with the ada-002 embedding, which has 1,536 dimensions.

All operations on the API are sized (and priced) in “tokens,” so it’s a good idea to employ the tiktoken counter, and keep an eye on your token limits. My 1,200-character chunks run around 220 tokens.

LangChain’s recursive text splitter does its best to honor sentences and paragraphs, but not semantics. It still feels like taking a favorite novel and chopping it up in a Cuisinart. Greg Kamradt has invented a semantic text splitter, which can detect and split based on topic changes, but its implementation in LangChain isn’t great.

Chroma’s query method takes a string, vectorizes it, and then does a similarity search against the embeddings in the database. Normally, you look to maximize “cosine similarity” but, with Chroma, you must minimize “cosine distance.” You can also call the OpenAI embedding function on your own, and then search by vector directly. Just be sure to use the same embedding model in all cases.

That covers the basics and the single-novel case. Next week, we’ll use RAG to compare two novels.

Sidebar: While writing this, I felt the need to review some points from War and Peace, so I pulled the book off the shelf … and then remembered I had just built a searchable database. Old habits die hard.