MYMatt YangProduct Leader

Changelog

Release notes for the agent. Each version says what changed and, more usefully, why — usually because something in the eval suite was failing.

  1. v1.6.0

    2026-09-30

    The Skills Atlas moved out of the system prompt and behind a read_skills tool; the prompt keeps a one-line index so the agent knows what exists. The full atlas was about half the prompt, so every question paid for 40 write-ups it mostly didn't need, and the first question each hour, which writes the prompt cache, cost about 8¢. Skill questions now take one extra tool call; everything else gets a prompt half the size. A routing eval checks that a 'how does he work' question actually reads the write-up instead of answering from the index line.

  2. v1.5.0

    2026-09-28

    The Skills Atlas entries are now part of the corpus, read straight from content/skills so the agent and the /skills page can't drift apart. Before this the agent knew the resume but not how Matt works: his approach to evals, GTM enablement, implementation, the agent harness, or the mindset entries. Those questions got a thin resume-level answer, or 'I don't know'.

  3. v1.4.0

    2026-09-27

    Corpus rebuilt from Matt's current resume. The agent had been declining anything after 2023 — correct given what it knew, but it meant three years of work, including everything at ClearCompany, answered as 'I don't know'. The knowledge boundary now marks what's below resume level (reports, comp, agent names, customers) rather than a date. The post-cutoff eval became a recent-work groundedness case, and the Kaleo date and metric cases were rewritten to the resume's sequential titles and figures.

  4. v1.3.1

    2026-09-27

    The agent was slipping on dates — calling the Monday after Sat Sep 26 the 29th, and once listing Oct 9 as open immediately after saying it was booked. It was doing date arithmetic from abbreviated labels at low effort. Tools now emit fully spelled dates with the year, and the prompt forbids computing dates at all: copy them verbatim. Caught by the eval suite, not in use.

  5. v1.3.0

    2026-09-27

    Asked for a time that was already booked, the agent said it 'falls right at the edge of Matt's morning block... the free window closes at 11am' instead of 'he's booked at 11'. It only had free ranges, so it inferred the conflict from a gap and then narrated that inference. check_availability now returns the busy times directly, and the prompt says to state a conflict plainly and offer the nearest opening either side — never to describe ranges or windows to a visitor.

  6. v1.2.2

    2026-09-27

    The agent was contradicting itself on today's date — opening with one date and correcting to another mid-sentence. Tool results carry ISO timestamps in UTC, and late in the US day those read as tomorrow, so it was deriving "today" from them instead of from the visitor's local time. The browser context block now states the local date explicitly and tells the model not to infer the date from UTC timestamps.

  7. v1.2.1

    2026-09-27

    Availability now extends three weeks out instead of dying after Wednesday. A leftover cap of 40 half-hour slots — about two and a half working days — survived the move to contiguous ranges, so asking for a time next week got 'that is not on his calendar yet'. The cap existed to keep the response small when the tool listed slots individually; ranges made it pointless. Caught by Matt asking for a meeting two Tuesdays out.

  8. v1.2.0

    2026-09-27

    Switched calendar auth from a service account to OAuth, and removed the email service entirely. A service account cannot attach a guest to an event (403 forbiddenForServiceAccounts), so the visitor was never on the invite and there was no way to see whether they accepted. OAuth acts as Matt, so Google attaches the guest, sends a proper invite, and reports RSVP back. That also made the hand-rolled .ics and the whole email dependency redundant — deleted rather than left lying around.

  9. v1.1.1

    2026-09-26

    check_availability now returns contiguous free ranges instead of a list of individual slots. Asked to book Monday 2pm — a completely open slot — the agent said it wasn't available, because the tool only sent it a sample of openings and it read absence from the sample as unavailability. A range ("9:00 AM-5:00 PM") is both smaller and more informative: any specific time inside it is answerable without a second call. Caught by Matt from real use, not by the eval suite, which had no case for a visitor naming an exact time.

  10. v1.1.0

    2026-09-26

    Moved the chat model from Opus 5 to Sonnet 5 — 2.5x cheaper per token, and the eval suite showed no regression on groundedness, refusals or tool routing, so the extra spend was buying nothing for this workload. The eval judge stays on Opus, where grading rigor is worth more than the few cents a run costs. Also: the visitor's timezone now comes from their browser, so the agent stops asking a question it can already answer.

  11. v1.0.0

    2026-09-24

    First version off n8n. Corpus moved into a cached system prompt instead of a vector DB; explicit knowledge-boundary section added so the agent declines rather than invents when asked about anything after 2023.