Part 2 of the “Raising a Lobster” series. Part 1 covered week one. This one covers the month of operations that followed.
Part 1 ended with JJ delivering a stock brief on time, every number filled in. I thought it was stable.
I spent most of March as its repairman.
Market updates at midnight
Starting the evening of February 27, JJ sent me a market update roughly every two hours, including in the middle of the night. That job was supposed to run once a day, after the market closed.
The next day I asked it directly:
It said there were none. The messages kept coming.
There were two causes. First, scheduled jobs have a wakeMode setting that defaults to now: after a container restart, every missed job runs again. Each time Zeabur restarted the container, I got another round. Second, every job run left behind a session that never got cleaned up, a bug in that version. They piled up. The first time I cleared them, there were 38.
After switching wakeMode to skip and clearing the sessions, things were quiet for two days, and then there were 36 again. Worse, one time after a cleanup and reload, a session file stayed locked and JJ went completely silent, answering everything with “All models failed.” In the end I combined “clear sessions, delete lock files, reload” into one command, saved it in my notes, and pasted it whenever the symptoms showed up.
The lobster starts making up stock prices
On the morning of March 5, the numbers in the stock brief looked off. The news it attached didn’t match that day’s market at all, and the “source links” went nowhere. JJ also slipped into Simplified Chinese now and then.
It turned out Yahoo Finance was blocking Zeabur’s IP addresses. When JJ couldn’t get data, it made up the numbers, the news and the links. It wasn’t trying to lie. It was trained to produce an answer.
I added three absolute rules to the top of SOUL.md, its persona file:
- Always reply in Traditional Chinese, whatever model is running.
- If data can’t be fetched, report an error. Never estimate or invent numbers, headlines or sources.
- No real URL, no “source link.”
I also changed the data sources: Taiwan’s stock exchange official API for local stocks, and Alpha Vantage for US stocks. On test day I used up the free US quota, and JJ reported “lookup failed (daily limit reached)” instead of guessing. It was the first time an AI saying “I couldn’t find it” made me happy.
The rules were never finished in one go. On March 10 I asked JJ about the World Baseball Classic. It stitched games from different days together and invented the scores. SOUL.md got another rule: if you can’t find the score, say so.
Upgrading: npm installs that vanish
Upgrades were rough from the start. On February 28, the first upgrade added a new security setting; until it was set, the gateway wouldn’t start, and the log only said it would retry in an hour.
On March 6 I tried to move to 2026.3.2. Following an online guide, I ran npm install inside the container, and the version number didn’t budge. I installed to another path, the container restarted halfway through, and I was back at the start. At one point I simply typed: “What is going on?”
The answer was simple. On Zeabur, anything installed inside the container disappears on restart. To upgrade, change the version tag of the Docker image and click save. An hour of my time, one line of knowledge.
Upgrades also had two side effects I learned to check every time:
- The Brave Search API key environment variable disappeared, so JJ couldn’t search and its answers went hollow.
- A setting that controls proactive direct messages reset to its default and had to be changed back by hand.
A better model, a smarter lobster
On March 7 I wanted JJ to compare prices: give it a product, and it lists prices from Taiwan’s online stores, cheapest first. I wrote a SKILL.md that called a Google Shopping search API.
On gpt-4o-mini, it ignored the steps entirely, did a quick web search, and told me the official price “starts at NT$19,900.” On Claude Sonnet, it followed the SKILL.md, called the API, and sorted the results from low to high. My reaction at the time: “Everyone was right. When you raise a lobster, don’t skimp on the model.”
I also tried the cheaper Claude Haiku, which handled it fine, so Haiku became the main model.
Two days later, half the US stock quotes failed again because Alpha Vantage allows only five requests a minute. I asked JJ to fix it. It changed the US lookups to run one at a time, 15 seconds apart, and documented the change. I typed: “I think JJ got smarter 🤣“
Three lobsters, finally a team
March was also when the three agents started working as a team.
- Delegation: JJ sent the advisor to compare iPhone prices, and the advisor handed the results back. I only talked to JJ.
- The advisor’s persona: I read that instead of writing an agent’s persona yourself, you should let the agent interview you. So JJ asked me ten questions and wrote a 163-line persona file for the advisor. The advisor’s catchphrase: “Wait, there’s an assumption here nobody has stated.” I gave it a hard question from work, and it didn’t offer me a single comforting word.
- Alfred the butler: The butler got connected to Google Calendar. Google had retired the old authorization flow, so I ended up running a tiny PowerShell server on my own PC to catch the authorization code. After that, I sent it a photo of a calendar, and it read it and created the events and reminders on its own. Its persona is Batman’s Alfred: holds the fort, reminds you before you ask, keeps it short.
A false alarm
On the morning of March 22, JJ pushed a warning:
The config file was fine. The daily backup saved it with a git commit. When the file hadn’t changed that day, git said there was nothing to commit, the upload step after it never ran, and JJ concluded the backup had failed. Adding the --allow-empty flag fixed it.
A classic bug: the system wasn’t broken. The logic that decides whether it’s broken was.
I started writing handoff notes
The biggest lesson of the month had nothing to do with the lobster.
Every debugging session meant opening a new Claude chat and re-explaining the config, the logs and everything I had already tried, and I often hit my usage limit halfway through. From March 5 I changed my approach. At the end of each chat, I asked Claude for a short summary of what we concluded, what comes next and what the next chat needs to know, and appended it to a session-log.md. Every new chat started with: “Please read session-log.md first.”
That log ran through March 22, and it became the skeleton of this post. I still work the same way with Claude Code on every project: update the progress file before stopping, and pick up from there next time.
A month later
By the end of March, JJ delivered its brief on time every morning, no longer made up numbers, and the three lobsters had clear roles. Looking back, though, I spent far more time repairing it that month than it saved me.
Next time: April to September, and why I shut the lobster down.