X Content Generator

A daily pipeline that researches AI news, drafts candidate posts, and scores them on five axes.

  • 2026–present
  • Personal

Why there is no measurement

The feedback loop has run for a month and stored nothing

No measurement

Why not

The measurement this project was built to produce is whether scored drafts correlate with real engagement. It cannot be produced, and the reason is the finding. Every run persists its candidates to a Supabase table, and every one of those writes has failed since the first run on 2026-07-21 — the table has row-level security enabled and no policy granting insert, so Postgres rejects the row. Twenty-five successful runs; zero rows stored. The write is wrapped in a try/except that prints and continues, because a save failure was not supposed to stop the digest going out, so the job exits 0 and the Telegram message arrives looking exactly as it should. I found this by querying the table during an audit of this page, not from using the system daily for a month. Everything downstream of that write — the fourteen-day repeat suppression, the engagement feedback, the whole self-improvement premise — has been running against an empty set the entire time.

Problem

Posting consistently about a fast-moving field means either writing from whatever you happened to read, or not posting. I wanted the drafting to start from what actually happened that day, and I wanted the choice of what to post to be a ranking against stated criteria rather than whichever draft I liked at the time.

Approach

Two model calls with a scoring step between them. The first researches the day's AI news and over-generates a pool of candidate posts across four formats, each required to cite a specific source. The second scores every candidate 0–10 on five axes and the pool is ranked, with the top few sent to Telegram for me to accept, edit or ignore. Every candidate was meant to be persisted so that logged engagement could feed back into later runs. That last part is where this project has something to say.

Components

  1. Trend research

    Evidence: UNVERIFIED

    Four fixed queries covering release news, research, funding and general AI news, each restricted to the last two days, so a run's candidates are not all angles on one story.

  2. Candidate generation

    Evidence: UNVERIFIED

    One model call drafting nine posts across four formats, each required to cite a source URL from the research and forbidden from asserting developments not present in it.

  3. Deliberate over-generation

    Evidence: UNVERIFIED

    Nine drafted to send five. Sampling is at temperature zero, so the spread the ranking needs has to come from asking for variety explicitly rather than from the sampler.

  4. Prompt-injection boundary

    Evidence: UNVERIFIED

    Both model calls treat their input as untrusted — the first because it is fetched from the web, the second because its input was written by the first from that web content.

  5. Five-axis scoring

    Evidence: UNVERIFIED

    Scores trend relevance, originality, engagement potential, voice fit and clarity, 0–10 each. Missing or malformed values are clamped to the range rather than trusted.

  6. Ranking

    Evidence: UNVERIFIED

    Sorts by total and takes the top five. No promotion gate, unlike the screener it shares a repo with — there is always a best five, and nothing worth blocking on a threshold.

  7. Candidate persistence

    Evidence: UNVERIFIED

    Writes every candidate, sent or not, to a Supabase table so that later runs can see what was already suggested and what performed.

    Fails on every run, and the run reports success

    No measurement

    Why not: Row-level security is enabled on the table with no insert policy, so every write is rejected with a 401. The call is wrapped so that a storage failure cannot stop the day's message going out, which is a defensible choice that turned a broken dependency into a log line. This is the component every other unverified marker on this page should be read against: the code is written, it is deployed, it runs daily, and it has never once done its job.

  8. Repeat suppression

    Evidence: UNVERIFIED

    Passes the last fourteen days of suggestions into the drafting prompt so the same angle is not proposed twice. Fourteen days, not the screener's ninety: AI news turns over fast.

    Inert, because it reads the table nothing is written to

    No measurement

    Why not: The logic is sound and has unit tests. It queries the same table the write above fails on, so it has returned an empty list on every run the system has ever done, and the drafting prompt has never once been told what it already suggested. Nothing in the output looks wrong. That is the problem with a silent dependency — the feature does not error, it simply never happens.

  9. Manual engagement logging

    Evidence: UNVERIFIED

    A small command-line tool for recording likes, replies and reposts against a sent idea. There is no API access to the account, so this step is a person typing numbers in.

  10. Weighted feedback

    Evidence: UNVERIFIED

    Ranks past ideas by likes plus twice replies plus three times reposts, and feeds the top few into the next drafting prompt as what has worked.

  11. Telegram delivery

    Evidence: UNVERIFIED

    Sends the ranked shortlist, retrying on rate limit, falling back to plain text if the markdown is rejected, and truncating at a line break rather than mid-token.

  12. Cost metering

    Evidence: UNVERIFIED

    Reports tokens, per-provider call counts and a priced total at the end of every run, which is how the cost of a run is known rather than assumed.

Two calls and a ranking

The generator does not ask a model for a good post. It asks for nine, then asks a second call to score them, then takes the top five. The split is the only structural idea in the project and it is worth stating plainly: generation and evaluation are different tasks, and a model asked to do both at once will justify whatever it produced.

Sampling runs at temperature zero, which removes the usual source of variety. The spread the ranking depends on therefore has to be requested — four named formats, an explicit instruction that candidates be distinct — rather than harvested from the sampler. Over-generating and discarding is the cheap way to get a ranking that means something.

Every candidate must cite a URL from that day’s research, and the prompt forbids asserting developments the research does not contain. That is the same instinct as the rest of this site: the interesting failure is not a model that refuses, it is a model that produces a confident, well-formed post about something that did not happen.

Untrusted input, twice

The research step fetches arbitrary web pages, so the drafting call treats its input as data rather than instruction. The scoring call does the same, and that one is easier to miss: its input is not the web, it is text another model call wrote from the web. A page that talks the drafter into a particular angle would otherwise get a second, unguarded shot at talking the scorer into a high mark for it.

What the audit found

This page was written by reading the code. Partway through, I checked whether the table the system writes to had anything in it. It had nothing.

The generator has run every morning since 2026-07-21. Twenty-five of those runs succeeded, the Telegram message arrived each time, and not one candidate was ever stored. The table has row-level security enabled and no policy permitting insert, so every write comes back rejected. The write is wrapped in a try/except that prints a warning and continues — deliberately, so that a storage problem could not stop the day’s message — and the effect is that the job exits zero with a failure in the middle of it.

Everything built on top of that write has been running against an empty set. The fourteen-day repeat suppression queries a table with no rows. The engagement feedback ranks a list that is always empty. The self-improvement loop, which is the entire reason the persistence exists, has never had a single row to improve from. None of this is visible in the output. The messages look right, because the part that produces them works.

Why this is on the site rather than fixed first

It will be fixed. But the sequence matters more than the fix, and writing the page first is the honest order.

I used this system daily for a month and did not notice. I would have described it, accurately as I understood it, as having a working feedback loop — and that description would have been wrong in the specific way this site exists to argue against: a claim about a system, made from the design rather than from the running thing, by the person best placed to know better.

What caught it was not a test. The unit tests pass; they cover the pure functions, and the failure is in a call that was designed not to raise. It was caught by one query against production, run because a page was going to make a claim and the claim needed checking. That is the whole argument of this site, demonstrated at my own expense, which is the only way a demonstration like this is worth anything.

The generalisable lesson is about the try/except, not about row-level security. Swallowing an error to protect a user-facing path is a reasonable instinct, and it silently converted a hard dependency into an optional one. If a failure is allowed to be non-fatal, something still has to be watching it — otherwise “non-fatal” quietly becomes “unobserved”, and the only person who finds out is whoever eventually goes looking.