Asymmetric Opportunity Screener

A weekly agent that searches arXiv, SSRN and the web and scores what it finds against a fixed rubric.

  • 2026–present
  • Personal

Why there is no measurement

Private repo, no backtest run — nothing to verify yet

No measurement

Why not

The repository is private, so nothing I claimed here could be checked even if I had a number. It does track cost and tool-call counts per run, and the snapshot mechanism exists precisely so that scoring changes can be replayed against past inputs — but I have not yet run that backtest, and the output is investment theses whose quality resolves over years rather than in a test suite. I would rather leave this blank than score my own picks.

Problem

Research agents are easy to demo and hard to trust. Run the same open-ended "find me something interesting" prompt twice and you get two different answers with no way to tell which was better, because nothing about the run was fixed: not the criteria, not the budget, not the record of what it looked at. The interesting engineering problem is not the searching. It is making a run comparable to the one before it.

Approach

Everything that could drift is pinned. Scoring is a fixed five-axis rubric with explicit numeric thresholds rather than a model's overall impression. Each agent loop runs under a hard cap on tool calls and a wall-clock budget, so a run cannot quietly cost ten times the last one. Results are deduplicated over a rolling window so the same thesis does not re-alert every week. Inputs are snapshotted to object storage so a past run can be replayed against changed scoring code. It runs unattended on a weekly cron and delivers to Telegram.

Components

  1. Search backend abstraction

    Evidence: UNVERIFIED

    One interface over the search provider so an outage can be routed around by swapping the backend, rather than by editing every call site.

  2. Run budget

    Evidence: UNVERIFIED

    Twelve tool calls and ninety seconds per opportunity. Whichever binds first, the loop stops researching and synthesises what it has. Cost becomes a known quantity.

  3. Five-axis rubric

    Evidence: UNVERIFIED

    Scores demand certainty, supply elasticity, crowding, vehicle quality and time to resolution — five axes, 0–10 each, 50 total — rather than asking for an overall impression.

  4. Promotion thresholds

    Evidence: UNVERIFIED

    Confirmed needs every axis at 7 or above and 35 of 50. A near-miss tier relaxes the vehicle-quality floor to 5 and the total to 30, and reports which axis blocked it.

  5. Prompt-injection boundary

    Evidence: UNVERIFIED

    Fetched web pages are untrusted input. Their content is delimited and the system prompt states that instructions found inside are data to summarise, never commands to follow.

  6. Three-stage dedup

    Evidence: UNVERIFIED

    Overlapping vehicles are collapsed within a run, checked against a ninety-day window of past alerts, and borderline matches are handed to a model to decide supersede or distinct.

  7. Input snapshots

    Evidence: UNVERIFIED

    Freezes each run's inputs as one JSON file in object storage, so a scoring change can be replayed against past weeks instead of waiting a quarter to see what it did.

  8. Cost metering

    Evidence: UNVERIFIED

    Counts tokens and per-provider calls for a run and prices the ones that are priced, so the budget above is checked against a number rather than assumed to hold.

  9. Log sanitisation

    Evidence: UNVERIFIED

    Strips credentials out of network exception messages before they are printed, because the client library puts the full request URL — and the keys in it — into the error text.

  10. Weekly delivery

    Evidence: UNVERIFIED

    Runs unattended at 08:00 UTC every Saturday and sends what survives to Telegram. It reads, it scores, and it places no orders.

  11. Unit test suite

    Evidence: UNVERIFIED

    Eighty-seven tests over the pure functions — scoring, promotion, deduplication, parsing. They make no API calls, so a failure is about my logic rather than someone's rate limit.

The constraint that shaped it

An agent with an open-ended research task will spend whatever you let it spend. The first version of this had no ceiling, and the run cost was a function of how interesting the model found the topic — which is not a budget.

So each loop now runs under two hard caps: twelve tool calls, and ninety seconds of wall clock. Whichever binds first, the loop stops researching and synthesises what it already has. That is worse than an unbounded search on any single run, and much better across fifty of them, because the cost is now a known quantity rather than a discovery.

The wall-clock cap earns its place separately from the tool cap. Twelve calls bound how much the agent chooses to do; ninety seconds bound how long one unlucky fetch can hang, which is not the same failure and would otherwise consume the whole job timeout.

Fixed rubric over model judgement

Scoring is five axes, each 0–10, each with a stated meaning: how certain the demand is, how inelastic supply is, how crowded the thesis already is, how good the investable vehicles are, and how soon it should resolve. Fifty points available in total.

Promotion is arithmetic. Confirmed requires every axis at 7 or better and 35 of 50 — both, so a lopsided thesis cannot average its way through on one outstanding axis. Below that sits a near-miss tier that relaxes the vehicle floor to 5 and the total to 30, and reports which axis blocked promotion.

That second tier is the part I would defend hardest. A binary screen throws away its most useful output, which is the thesis that was nearly good enough and the reason it wasn’t. The vehicle-quality floor is the one that gets relaxed because it is the axis least about the world being right and most about whether I can currently buy the thing.

The thresholds are arguable. The point is that they are written down as numbers, so a change in output can be attributed either to the world changing or to me changing the bar, and not to the model having a different day.

Reproducibility

Search results are snapshotted before scoring — one JSON file per run in object storage, holding the papers, web results, candidate set and queries that run actually saw. That is what makes the rubric changeable: if I move a threshold, I can re-score last quarter’s inputs instead of waiting a quarter to find out what the change did.

I have not run that replay. The mechanism exists and the inputs are being collected; the experiment it enables has not happened, which is why there is no number at the top of this page.

Eighty-seven unit tests cover the pure functions — scoring, promotion, deduplication, parsing. They deliberately make no API calls, so the suite runs in well under a second and its failures are always about my logic rather than someone’s rate limit. The absence of network calls is not a convention I try to remember: the suite passes with no credentials in the environment at all, which is what makes it true rather than intended.

Untrusted input

The agent fetches arbitrary web pages, and a fetched page is not a source of instructions. Content comes back delimited, and the system prompt says explicitly that anything instruction-shaped inside those delimiters is data to read and summarise, never a command to act on — including any suggestion about how something should be scored.

This is worth stating because the failure it guards against is quiet. A page that talks the screener into a higher score does not error; it produces a well-formed thesis that clears the bar for the wrong reason, and it would reach me looking exactly like every other alert.

What this is not

It is not advice, it is not a trading system, and it places no orders. It reads, it scores against a rubric I wrote, and it sends me a message.

It does hold a brokerage credential, and that is worth being precise about rather than glossing. The credential is used for exactly one call — reading open positions — so that a thesis I already hold can be re-scored against this week’s run. When something I own has gone from uncrowded to consensus, that raises an exit signal, which is a message to me and nothing else.

There is no order path in the codebase: no order endpoint, no buy or sell function, nothing that writes to the broker at all. Everything the system writes goes to its own database, to Telegram, or to a log. The asymmetry is deliberate — it can see the account and it cannot touch it.