Every week, 27 sources dump around 300 articles on me to sort through. A pipeline brings them down to a 10-minute read, on its own, with an AI in the middle, without me opening my terminal.

One Saturday, it crashed. Not because of the AI, but because of the trust I had placed in it.

That's what I want to share here. The way I see it, an engineer's job isn't just to write code: it's to understand a problem and solve it. Whether an AI holds the keyboard or not changes nothing. That Saturday, if the pipeline got away from me, it's because I hadn't understood the problem well enough to see it coming.

Why a tech watch

I built it for two reasons.

The first: to have a personal tech watch, tailored to what I want to follow given where I want to go. Cloud, DevOps and AI move every month, and I didn't want yet another generic feed. I wanted to follow what matters for my path.

The second: to experiment with real automation, end to end. Something that runs on its own, in production, with everything that implies.

The concrete problem is volume. 27 feeds means around 300 articles a week. Reading them all is impossible, ignoring them would be worse. What's missing isn't content, it's curation.

That's where the AI comes in, and only there. Not because it's fashionable, but because it handles the part that simple rules can't: reading the 300 articles, keeping the relevant ones, removing duplicates, summarizing them. The rest of the pipeline, fetching the feeds and publishing, is plain Python.

Concretely, every Saturday, a GitHub Actions cron fetches the articles from the last 7 days, Claude Haiku selects and summarizes the most relevant ones in French and English, and the pipeline publishes an edition on my portfolio. Three Python scripts, a few cents a week. It has been running for four months, 19 editions so far. The links are at the end of the article if you want them.

The AI's place in the project

I set one rule from the start: the AI does the work, but nothing gets published without me having read it.

The pipeline puts nothing online on its own. It prepares the edition and opens a Pull Request. That review serves two purposes. First, doing my tech watch: re-reading the edition is already reading it, and that's the whole point of the project. Second, checking the AI's work, that the summaries match the articles and that no duplicate slipped through. For each article, the original source stays accessible, if I want to dig deeper.

Upstream, same logic: the instruction is written into the prompt, factual summaries, no opinion or commentary. I don't want the AI interpreting the news for me, only giving me the fact and what it means.

The crash

Two weeks after going live, on a Saturday morning, the CI alerts me: the workflow fails at the summarize step. No edition that day.

In itself, the incident is minor. The CI caught it right away, no data was lost, and it took me half an hour to fix. But the cause was worth pausing on, because it's a trap.

My prompt asked Claude to summarize the thirty-odd selected articles in one go, each with a translated title, two summaries and a category, all in a single JSON response. Two instructions were written there in plain sight:

Summarize ALL articles without exception.
Respond with ONLY valid JSON.

That week, the selection was dense. The response exceeded the token budget I had set, 8192. Claude wrote up to the limit, then stopped, mid-JSON. My script got an object cut clean off, impossible to read.

The problem didn't come from Claude, but from me. I had taken a prompt instruction for a guarantee. Writing "respond with valid JSON" doesn't limit the length of the response, and "summarize everything without exception" pushes it right past the limit. A prompt expresses an intention, it doesn't guarantee a result.

But it got worse. My parser was hiding the problem: when it didn't find usable JSON, instead of raising an error, it started over from the beginning of the text with a default value.

# The buggy version: default=0 hides the missing JSON
start = min((text.find(c) for c in '[{' if text.find(c) != -1), default=0)

That behavior hid the problem instead of flagging it. The day the response was cleanly truncated, everything gave way at once, with no lead to trace the cause.

What I took away

An AI's output is unreliable data, and should be treated as such. We know to validate user input, handle a network timeout, be wary of a third-party API. A model's response calls for the same caution: it isn't deterministic, and nothing guarantees it fits within the space you planned. I validate it and parse it defensively; I no longer treat it as a safe value coming from my own code.

An error should be visible. My first fix wasn't to make the pipeline smarter, but to make it fail out loud. If the response contains no usable JSON, the script now raises an explicit error, with the start of the text it received. A visible bug gets fixed in minutes; a silent one can cost you days.

candidates = [text.find(c) for c in '[{' if text.find(c) != -1]
if not candidates:
    raise ValueError(f"No JSON found in response: {text[:200]}")
start = min(candidates)

The real problem was in the design, not the setting. I raised max_tokens to 16384 and added 3 attempts. It helps, but it only pushes the ceiling back: the day the radar keeps more articles or longer summaries, the same bug returns. The right fix, the one I'd make starting from scratch, isn't to enlarge the response, but to stop asking for one huge one. Split into batches of 5 or 10 articles, handle one article per call, or enforce an output schema. Bound what you ask the model for instead of hoping it fits.

for attempt in range(3):
    response = client.messages.create(
        model="claude-haiku-4-5-20251001",
        max_tokens=16384,          # was: 8192
        messages=[{"role": "user", "content": prompt}])
    try:
        return extract_json(response.content[0].text)
    except (json.JSONDecodeError, ValueError) as e:
        print(f"Attempt {attempt + 1}/3 failed: {e}")
        if attempt == 2:
            raise

Understanding the problem stays my job, whether the AI writes the code or not. The AI saved me a lot of time on the tedious part, sorting 300 articles a week. But the day I trusted it without checking, it let me down without warning. The way I see it, an engineer isn't only there to write code, they're there to understand a problem and solve it. Whether an AI or I write it changes nothing: if I don't understand what it produces, or where it can break, I can't see it coming. That's exactly what happened that Saturday.

And now

The radar is still running. Every Saturday, without me, it brings 300 articles down to a 10-minute read. I went from a tech watch I did when I thought about it, meaning rarely, to a regular one I actually read.

But I don't leave it on its own. The AI does the heavy lifting, I keep my hand on what matters: what gets published, what's accurate, and what breaks. That's the lesson I'm keeping for the next project where I put an AI in production. It's an excellent tool, as long as you stay the one who understands what it does.

If this interests you, this week's edition is at arthurbernard.dev/en/tech-radar, and the code is public at github.com/TuroTheReal/weekly-tech-radar.