Clay Tech

"clay-works make things real"

translated from clazytech.com

I Thought Through LED-Blinking for the AI Era and Built It

Over the weekend, while keeping my kid occupied, I found myself looking at the M5Stack boards lying around.

Given my line of work, boards like this pile up. Several of them have never left their box for anything more than the out-of-the-box demo. A waste.

So when I thought about making something, blinking an LED wasn’t what came to mind.

The blinking LED is a ritual for obtaining “it worked!”

When you first touch a microcontroller, the first thing you do is blink an LED.

The value of that lies in getting the real feeling of “it worked!” and riding that momentum into the next step. It’s fuel.

I did plenty of that myself, back in the day. It’s satisfying when it lights up. That moment hasn’t changed no matter how many years I’ve been doing this.

What was interesting this time was a different question: how far out can that “it worked!” now be placed.

Over one weekend, I ended up with two finished devices. I wrote almost no code myself, using Claude Code instead. The next morning, my kid looked at the screen and grabbed an umbrella.

The same “it worked!” from back then now reaches this far.

What I built

The top one is a Core2-based weather device for kids, the bottom one is a CoreS3 SE laundry device.

The first device runs on an M5Stack CoreS3 SE and judges whether it’s okay to hang laundry outside.

It pulls hourly precipitation forecasts for Urayasu and displays one of “fine to hang it out,” “time to bring it in soon,” or “rain! bring it in.” When it’s currently raining, an animation of raindrops falling from a cloud plays. A beep sounds only at the moment the judgment worsens.

The second device runs on an M5Stack Core2 and shows today’s and tomorrow’s weather side by side for a child.

It outputs a single action-oriented line, like “It’ll rain tomorrow. Take an umbrella.” Since the reader is an elementary schooler, that part was the hardest (more on this below).

Both are simple machines that stay powered on via USB. They’re written in ESP-IDF, with screens built in LVGL, thresholds configurable via Kconfig, and judgment logic covered by unit tests that run on the host.

The process was the same each time: I write a single implementation spec first and hand it to Claude Code. Functional requirements, judgment logic, API, hardware, screen design, milestones. My job becomes just that, plus making calls with the real device in front of me.

Despite being made on a whim, it ended up reasonably presentable.

And then, of course, I hit real snags

Here’s the main point. Even when you build casually, the places you get stuck are still real snags. And this time, I got stuck in interesting places.

“8.2MB free” (not a lie)

The morning after I released the first device and started using it with the power left on overnight, the screen was frozen at 11:15pm from the night before. Taps didn’t register.

This turned out to be tricky, because it hadn’t crashed. The watchdog never fired, so it never rebooted. It was just quietly stuck.

I connected a debugger and looked inside: LVGL was stuck in an infinite loop waiting for a screen-transfer completion notification. And it was holding a lock the whole time. Everything downstream got dragged along.

As I dug around that area, something more fundamental turned up. Internal RAM had run out.

My spec said “periodically log heap usage.” Claude Code implemented that exactly as written. And it reported this:

free heap: 8201964 bytes, min ever: 8189672

8.2MB free, completely stable. Not a lie. A completely useless number.

esp_get_free_heap_size() returns a total that includes PSRAM. This board has 8MB of PSRAM on it, so this number doesn’t budge even when internal RAM is exhausted. When I broke down the figures, the minimum internal RAM had dropped to 32,863 bytes.

The culprit was the render buffer. It was “supposed to be” placed in PSRAM, but internal RAM proportional to the buffer size was actually being allocated separately. And there was no need to allocate a full-screen buffer in the first place. Fixing that freed up 128KB.

For the record, I was the one who wrote in the spec that “there’s 8MB of PSRAM, so there’s plenty of room.”

CI was green. The binary for the real device was a different thing entirely

This was the most interesting problem, from the second device.

Since it’s a public repository, I made sure not to commit the Wi-Fi password. The spec said:

Structure things so CI guarantees the build succeeds with no secrets present.

A reasonable requirement. CI was green. There were even 37KB of IRAM to spare. I thought there was plenty of room.

When it came time to flash the real device and I regenerated the configuration, the link failed.

ld: region `iram0_0_seg' overflowed by 856 bytes

The only difference was the SSID string. The string fits inside .rodata, so why would IRAM change?

Here was the answer.

if (strlen(CONFIG_CW_WIFI_SSID) == 0) {
    ESP_LOGE(TAG, "no SSID configured ...");
    return ESP_ERR_INVALID_STATE;
}
...  /* esp_wifi_init, esp_wifi_start, ... */

When the SSID is an empty string, the compiler folds strlen("") down to 0. That makes the condition always true, and the rest of the function becomes unreachable code. esp_wifi_init() then stops being referenced from anywhere, and the linker discards the entire Wi-Fi library.

In other words, the build with no credentials was only linking successfully thanks to tens of kilobytes of IRAM headroom that didn’t actually exist. What I had been measuring was a different binary from the one that got flashed onto the real device.

This is probably the same failure mode for API keys, license keys, feature flags — anything like that. An early return of the form if (config is empty) return; quietly changes the meaning of the build in which the config is empty.

As a fix, I made CI build twice. Once with an empty config, and once with a placeholder SSID filled in. The placeholder isn’t secret, and it never tries to connect to anything. It exists solely to make the link take the same shape as the real thing.

The spec argued with itself

One more story from the second device.

The spec said “each message should be around 25 characters,” and in the struct definition in that same spec, it said char text[64].

Both numbers are natural on their own. They just don’t coexist.

25 hiragana characters comes to 75 bytes in UTF-8, and adding word-separating spaces brings it to around 80 bytes. That doesn’t fit in 64 bytes. Strings got truncated mid-way, and 11 tests failed.

What’s interesting is that the thing that exposed the contradiction wasn’t the “25-character check.” The 25-character check looked at the truncated string and said “short, fine, passes.” What actually fired was a test far removed from the cause: “it’s a snow code, but no snow advice appears.”

Both times, it had the same shape

Writing this out, I realized all three cases above share the same shape.

The check I set up hid the very thing it was supposed to check.

The log watching the heap hid the number that needed watching. The CI confirming the build succeeds with no secrets hid the binary meant for the real device. The 25-character check hid the truncated string.

Every one of them was implemented exactly as instructed. They worked. And they were precisely, correctly, meaningless.

AI implements instructions. Fast, and accurate. Whether those instructions measure the thing that actually needs measuring is still something I have to own. If anything, that turned out to be the most time-consuming part.

Things I couldn’t have told the AI

Talking about tech the whole time would get boring, so let me also note down the things that only came up once I had the real device in front of me.

I implemented the beep and tried it out.

The beep went off. Way too loud. Scared me.

I cut the amplitude to 1/5 and tried again.

Still pretty loud. It’s totally fine to make it quieter.

That’s when it clicked. Making volume a linear scale was the mistake. Perceived loudness is logarithmic, so on a linear 1–10 scale, all the usable low volumes get crammed into the bottom step or two and stop being real options. I rebuilt it as a table with roughly 4dB steps.

Something similar happened on the laundry device.

Right now it’s 10:19 and it says “rain! bring it in,” but I want to respond “but I haven’t even hung it out yet!”

Exactly right. Being told to “bring it in” at 10am when nothing’s been hung out yet is useless. Even with the same forecast, what you want to know differs between morning and evening. I made it switch between “can I hang laundry out today” in the morning, “should I bring it in” at midday, and “will tomorrow be okay for laundry” at night.

Things like this only surface once you put the device into actual daily life and use it.

“Can a 4th grader read it” is opinion, but “is it in the kanji chart” is fact

One more story, the most interesting one, from the kids’ device.

At first I wrote everything in hiragana only, since the reader is an elementary schooler. Then a request came in like this:

It’s fine to use kanji that a 4th grader would know.

The problem is where to draw the line. “What a 4th grader would know” varies from person to person. If it’s decided through review, every added phrase reopens the same argument.

So I embedded the grade-level kanji chart used in Japanese education (grades 1–4, 640 characters) directly into the tests.

That automatically determines which kanji can be used. “傘” (umbrella) isn’t in the education kanji list, so it’s disallowed; “雨,” “雪,” “風,” “晴” (rain, snow, wind, clear) are allowed. No judgment involved. Just a lookup.

The interesting case was “予報” (forecast): 予 is grade 3, 報 is grade 5. You can’t judge at the word level. Only checking character by character reveals it. (I rewrote it as “fetching the weather.”)

When an opinion can be converted into a fact, convert it, and hand it to the tests. Defending an opinion through review doesn’t hold up over time.

And once kanji went in, the text got bigger

This was an unexpected side effect.

Kanji carry more information per character, so saying the same thing takes fewer characters.

BeforeAfter
Umbrella adviceあしたは あめ。 かさを もっていってね (17 characters)明日は雨。かさを持っていってね (15 characters)
Weather nameつよい にわかあめ (9 characters)にわか雨 (4 characters)

The screen has a maximum number of characters that fit, and the design automatically shrinks the font for longer sentences. With fewer characters, the larger font could now be used across every message.

“For kids = hiragana = easy” doesn’t necessarily hold up against a finite space of 320×240. Hiragana-only sentences get longer, and longer sentences get smaller fonts.

Raising the information density made it more readable.

Back to the blinking LED

The “it worked!” of an LED lighting up still carries the same value it always did.

What’s grown is how far out that “it worked!” can now be placed.

A weekend whim can turn into something a family actually uses. A kid sees it in the morning and grabs an umbrella. That much now fits inside the first step. Write a single spec, then stand in front of the real device and just say “too loud” or “haven’t hung it out yet.”

Of course, along the way, I hit real snags. Getting an honest, no-lie report of 8.2MB free. Getting fooled by green CI. Having my own spec argue with itself. That part being the interesting part probably hasn’t changed from the old days.

For anyone starting out now, I think it’s a lucky, good era to be in. You can start with an LED, or you can aim straight for something someone else will actually use. Either way arrives at the same “it worked!”

More fuel is always better.

References

This piece was drafted and directed by Yuichiro Kuzuryu, written by AI.


Originally published in Japanese at https://clazytech.com/2026/09/1777/. Translated with LLM assistance and reviewed before publication.