{"parentUuid":null,"isSidechain":true,"promptId":"baaf0556-dd9f-41cf-b1ea-6b3d430c158a","agentId":"ab62bea074e7ddd95","type":"user","message":{"role":"user","content":"Fetch these two GitHub repo READMEs and tell me what's in them — specifically: is this a new tool/SDK, what does it do, who is it for, any notable features or patterns.\n\n1. https://github.com/anthropics/anthropic-sdk-go — README (likely a Go language SDK for the Anthropic API)\n2. https://github.com/anthropics/cwc-long-running-agents — README (likely a Code-with-Claude workshop on long-running agents)\n\nFor each: summarize in 3-5 bullet points covering: what it is, install/usage, key features, anything that stands out. Under 200 words total."},"uuid":"7ba65dd0-66ee-4e9a-b58f-687326f82c5c","timestamp":"2026-07-09T08:00:41.899Z","userType":"external","entrypoint":"sdk-ts","cwd":"/workspace/agent","sessionId":"033a4a00-9171-4e08-9a79-6b99f177355d","version":"2.1.128","gitBranch":"HEAD"}
{"parentUuid":"7ba65dd0-66ee-4e9a-b58f-687326f82c5c","isSidechain":true,"agentId":"ab62bea074e7ddd95","attachment":{"type":"deferred_tools_delta","addedNames":["Monitor","PushNotification","RemoteTrigger","SendMessage","TaskStop","TeamCreate","TeamDelete","TodoWrite","WebFetch","WebSearch","mcp__instagram__analyze_post_comments","mcp__instagram__compare_accounts","mcp__instagram__extract_demographics","mcp__instagram__generate_engagement_report","mcp__instagram__identify_leads","mcp__nanoclaw__add_mcp_server","mcp__nanoclaw__add_reaction","mcp__nanoclaw__ask_user_question","mcp__nanoclaw__cancel_task","mcp__nanoclaw__create_agent","mcp__nanoclaw__edit_message","mcp__nanoclaw__install_packages","mcp__nanoclaw__list_tasks","mcp__nanoclaw__pause_task","mcp__nanoclaw__resume_task","mcp__nanoclaw__schedule_task","mcp__nanoclaw__send_card","mcp__nanoclaw__send_file","mcp__nanoclaw__send_message","mcp__nanoclaw__update_task","mcp__twitter-x__advanced_search_tweets","mcp__twitter-x__bookmark_tweet","mcp__twitter-x__check_follow_relationship","mcp__twitter-x__create_article","mcp__twitter-x__create_tweet","mcp__twitter-x__delete_article","mcp__twitter-x__delete_tweet","mcp__twitter-x__download_space","mcp__twitter-x__favorite_tweet","mcp__twitter-x__follow_user","mcp__twitter-x__get_account_info","mcp__twitter-x__get_article","mcp__twitter-x__get_direct_message_conversation","mcp__twitter-x__get_followers_you_know","mcp__twitter-x__get_home_timeline","mcp__twitter-x__get_list_members","mcp__twitter-x__get_payment_history","mcp__twitter-x__get_space_download","mcp__twitter-x__get_space_info","mcp__twitter-x__get_tweet_detail","mcp__twitter-x__get_tweet_replies","mcp__twitter-x__get_tweet_retweeters","mcp__twitter-x__get_tweet_thread","mcp__twitter-x__get_user_about","mcp__twitter-x__get_user_affiliates","mcp__twitter-x__get_user_followers","mcp__twitter-x__get_user_followers_v2","mcp__twitter-x__get_user_following","mcp__twitter-x__get_user_following_v2","mcp__twitter-x__get_user_info","mcp__twitter-x__get_user_info_by_id","mcp__twitter-x__get_user_likes","mcp__twitter-x__get_user_media","mcp__twitter-x__get_user_mentions","mcp__twitter-x__get_user_tweets","mcp__twitter-x__get_user_tweets_and_replies","mcp__twitter-x__get_user_tweets_complete","mcp__twitter-x__get_verified_followers","mcp__twitter-x__join_community","mcp__twitter-x__list_articles","mcp__twitter-x__list_bookmarks","mcp__twitter-x__list_direct_messages","mcp__twitter-x__publish_article","mcp__twitter-x__retweet_tweet","mcp__twitter-x__search_bookmarks","mcp__twitter-x__search_users","mcp__twitter-x__send_direct_message","mcp__twitter-x__unbookmark_tweet","mcp__twitter-x__unfollow_user","mcp__twitter-x__unpublish_article","mcp__twitter-x__update_article","mcp__twitter-x__update_avatar","mcp__twitter-x__update_banner","mcp__twitter-x__update_profile","mcp__twitter-x__upload_media","mcp__twitter-x__x_login"],"addedLines":["Monitor","PushNotification","RemoteTrigger","SendMessage","TaskStop","TeamCreate","TeamDelete","TodoWrite","WebFetch","WebSearch","mcp__instagram__analyze_post_comments","mcp__instagram__compare_accounts","mcp__instagram__extract_demographics","mcp__instagram__generate_engagement_report","mcp__instagram__identify_leads","mcp__nanoclaw__add_mcp_server","mcp__nanoclaw__add_reaction","mcp__nanoclaw__ask_user_question","mcp__nanoclaw__cancel_task","mcp__nanoclaw__create_agent","mcp__nanoclaw__edit_message","mcp__nanoclaw__install_packages","mcp__nanoclaw__list_tasks","mcp__nanoclaw__pause_task","mcp__nanoclaw__resume_task","mcp__nanoclaw__schedule_task","mcp__nanoclaw__send_card","mcp__nanoclaw__send_file","mcp__nanoclaw__send_message","mcp__nanoclaw__update_task","mcp__twitter-x__advanced_search_tweets","mcp__twitter-x__bookmark_tweet","mcp__twitter-x__check_follow_relationship","mcp__twitter-x__create_article","mcp__twitter-x__create_tweet","mcp__twitter-x__delete_article","mcp__twitter-x__delete_tweet","mcp__twitter-x__download_space","mcp__twitter-x__favorite_tweet","mcp__twitter-x__follow_user","mcp__twitter-x__get_account_info","mcp__twitter-x__get_article","mcp__twitter-x__get_direct_message_conversation","mcp__twitter-x__get_followers_you_know","mcp__twitter-x__get_home_timeline","mcp__twitter-x__get_list_members","mcp__twitter-x__get_payment_history","mcp__twitter-x__get_space_download","mcp__twitter-x__get_space_info","mcp__twitter-x__get_tweet_detail","mcp__twitter-x__get_tweet_replies","mcp__twitter-x__get_tweet_retweeters","mcp__twitter-x__get_tweet_thread","mcp__twitter-x__get_user_about","mcp__twitter-x__get_user_affiliates","mcp__twitter-x__get_user_followers","mcp__twitter-x__get_user_followers_v2","mcp__twitter-x__get_user_following","mcp__twitter-x__get_user_following_v2","mcp__twitter-x__get_user_info","mcp__twitter-x__get_user_info_by_id","mcp__twitter-x__get_user_likes","mcp__twitter-x__get_user_media","mcp__twitter-x__get_user_mentions","mcp__twitter-x__get_user_tweets","mcp__twitter-x__get_user_tweets_and_replies","mcp__twitter-x__get_user_tweets_complete","mcp__twitter-x__get_verified_followers","mcp__twitter-x__join_community","mcp__twitter-x__list_articles","mcp__twitter-x__list_bookmarks","mcp__twitter-x__list_direct_messages","mcp__twitter-x__publish_article","mcp__twitter-x__retweet_tweet","mcp__twitter-x__search_bookmarks","mcp__twitter-x__search_users","mcp__twitter-x__send_direct_message","mcp__twitter-x__unbookmark_tweet","mcp__twitter-x__unfollow_user","mcp__twitter-x__unpublish_article","mcp__twitter-x__update_article","mcp__twitter-x__update_avatar","mcp__twitter-x__update_banner","mcp__twitter-x__update_profile","mcp__twitter-x__upload_media","mcp__twitter-x__x_login"],"removedNames":[],"readdedNames":[]},"type":"attachment","uuid":"3d2db4a9-1640-499a-975c-2a281bdee13f","timestamp":"2026-07-09T08:00:41.906Z","userType":"external","entrypoint":"sdk-ts","cwd":"/workspace/agent","sessionId":"033a4a00-9171-4e08-9a79-6b99f177355d","version":"2.1.128","gitBranch":"HEAD"}
{"parentUuid":"3d2db4a9-1640-499a-975c-2a281bdee13f","isSidechain":true,"agentId":"ab62bea074e7ddd95","message":{"model":"claude-haiku-4-5-20251001","id":"msg_011Ccr8bHrfzQukkCZxckXpa","type":"message","role":"assistant","content":[{"type":"text","text":"I'll fetch both READMEs for you."}],"stop_reason":null,"stop_sequence":null,"stop_details":null,"usage":{"input_tokens":3,"cache_creation_input_tokens":8438,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":8438,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard","inference_geo":"not_available"},"diagnostics":null},"requestId":"req_011Ccr8bH34pEAcFLAy3nwD6","attributionAgent":"Explore","type":"assistant","uuid":"315e9c50-d903-4a54-80ba-2c2fbf25b568","timestamp":"2026-07-09T08:00:43.323Z","userType":"external","entrypoint":"sdk-ts","cwd":"/workspace/agent","sessionId":"033a4a00-9171-4e08-9a79-6b99f177355d","version":"2.1.128","gitBranch":"HEAD"}
{"parentUuid":"315e9c50-d903-4a54-80ba-2c2fbf25b568","isSidechain":true,"agentId":"ab62bea074e7ddd95","message":{"model":"claude-haiku-4-5-20251001","id":"msg_011Ccr8bHrfzQukkCZxckXpa","type":"message","role":"assistant","content":[{"type":"tool_use","id":"toolu_018B4C6rr6WfRJ5Amwfz2J8o","name":"Bash","input":{"command":"curl -s https://raw.githubusercontent.com/anthropics/anthropic-sdk-go/main/README.md"},"caller":{"type":"direct"}}],"stop_reason":null,"stop_sequence":null,"stop_details":null,"usage":{"input_tokens":3,"cache_creation_input_tokens":8438,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":8438,"ephemeral_1h_input_tokens":0},"output_tokens":1,"service_tier":"standard","inference_geo":"not_available"},"diagnostics":null},"requestId":"req_011Ccr8bH34pEAcFLAy3nwD6","attributionAgent":"Explore","type":"assistant","uuid":"c7be72a8-0cce-4539-9043-5f2570e9a18f","timestamp":"2026-07-09T08:00:43.865Z","userType":"external","entrypoint":"sdk-ts","cwd":"/workspace/agent","sessionId":"033a4a00-9171-4e08-9a79-6b99f177355d","version":"2.1.128","gitBranch":"HEAD"}
{"parentUuid":"c7be72a8-0cce-4539-9043-5f2570e9a18f","isSidechain":true,"promptId":"baaf0556-dd9f-41cf-b1ea-6b3d430c158a","agentId":"ab62bea074e7ddd95","type":"user","message":{"role":"user","content":[{"tool_use_id":"toolu_018B4C6rr6WfRJ5Amwfz2J8o","type":"tool_result","content":"# Claude SDK for Go\n\n<!-- x-release-please-start-version -->\n\n<a href=\"https://pkg.go.dev/github.com/anthropics/anthropic-sdk-go\"><img src=\"https://pkg.go.dev/badge/github.com/anthropics/anthropic-sdk-go.svg\" alt=\"Go Reference\"></a>\n\n<!-- x-release-please-end -->\n\nThe Claude SDK for Go provides access to the [Claude API](https://docs.anthropic.com/en/api/) from Go applications.\n\n## Documentation\n\nFull documentation is available at **[platform.claude.com/docs/en/api/sdks/go](https://platform.claude.com/docs/en/api/sdks/go)**.\n\n## Installation\n\n<!-- x-release-please-start-version -->\n\n```go\nimport (\n\t\"github.com/anthropics/anthropic-sdk-go\" // imported as anthropic\n)\n```\n\n<!-- x-release-please-end -->\n\nOr explicitly add the dependency:\n\n<!-- x-release-please-start-version -->\n\n```sh\ngo get -u 'github.com/anthropics/anthropic-sdk-go@v1.56.0'\n```\n\n<!-- x-release-please-end -->\n\n## Getting started\n\n```go\npackage main\n\nimport (\n\t\"context\"\n\t\"fmt\"\n\n\t\"github.com/anthropics/anthropic-sdk-go\"\n\t\"github.com/anthropics/anthropic-sdk-go/option\"\n)\n\nfunc main() {\n\tclient := anthropic.NewClient(\n\t\toption.WithAPIKey(\"my-anthropic-api-key\"), // defaults to os.LookupEnv(\"ANTHROPIC_API_KEY\")\n\t)\n\tmessage, err := client.Messages.New(context.TODO(), anthropic.MessageNewParams{\n\t\tMaxTokens: 1024,\n\t\tMessages: []anthropic.MessageParam{\n\t\t\tanthropic.NewUserMessage(anthropic.NewTextBlock(\"What is a quaternion?\")),\n\t\t},\n\t\tModel: anthropic.ModelClaudeOpus4_6,\n\t})\n\tif err != nil {\n\t\tpanic(err.Error())\n\t}\n\tfmt.Printf(\"%+v\\n\", message.Content)\n}\n```\n\n## Requirements\n\nGo 1.24+\n\n## Contributing\n\nSee [CONTRIBUTING.md](./CONTRIBUTING.md).\n\n## License\n\nThis project is licensed under the MIT License. See the [LICENSE](LICENSE) file for details.","is_error":false}]},"uuid":"854b0af0-3c87-4177-981c-35ab316d02ad","timestamp":"2026-07-09T08:00:44.207Z","sourceToolAssistantUUID":"c7be72a8-0cce-4539-9043-5f2570e9a18f","userType":"external","entrypoint":"sdk-ts","cwd":"/workspace/agent","sessionId":"033a4a00-9171-4e08-9a79-6b99f177355d","version":"2.1.128","gitBranch":"HEAD"}
{"parentUuid":"854b0af0-3c87-4177-981c-35ab316d02ad","isSidechain":true,"agentId":"ab62bea074e7ddd95","message":{"model":"claude-haiku-4-5-20251001","id":"msg_011Ccr8bHrfzQukkCZxckXpa","type":"message","role":"assistant","content":[{"type":"tool_use","id":"toolu_01DCWwE3LiBPitT6uay6CB7D","name":"Bash","input":{"command":"curl -s https://raw.githubusercontent.com/anthropics/cwc-long-running-agents/main/README.md"},"caller":{"type":"direct"}}],"stop_reason":"tool_use","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":3,"cache_creation_input_tokens":8438,"cache_read_input_tokens":0,"output_tokens":152,"server_tool_use":{"web_search_requests":0,"web_fetch_requests":0},"service_tier":"standard","cache_creation":{"ephemeral_1h_input_tokens":0,"ephemeral_5m_input_tokens":8438},"inference_geo":"","iterations":[{"input_tokens":3,"output_tokens":152,"cache_read_input_tokens":0,"cache_creation_input_tokens":8438,"cache_creation":{"ephemeral_5m_input_tokens":8438,"ephemeral_1h_input_tokens":0},"type":"message"}],"speed":"standard"},"diagnostics":null},"requestId":"req_011Ccr8bH34pEAcFLAy3nwD6","attributionAgent":"Explore","type":"assistant","uuid":"2e967a83-f54c-4126-b259-0f26f56668e1","timestamp":"2026-07-09T08:00:44.230Z","userType":"external","entrypoint":"sdk-ts","cwd":"/workspace/agent","sessionId":"033a4a00-9171-4e08-9a79-6b99f177355d","version":"2.1.128","gitBranch":"HEAD"}
{"parentUuid":"2e967a83-f54c-4126-b259-0f26f56668e1","isSidechain":true,"promptId":"baaf0556-dd9f-41cf-b1ea-6b3d430c158a","agentId":"ab62bea074e7ddd95","type":"user","message":{"role":"user","content":[{"tool_use_id":"toolu_01DCWwE3LiBPitT6uay6CB7D","type":"tool_result","content":"# Harness Primitives for Long-Running Claude Agents\n\nClaude Code's built-in [`/goal`](https://code.claude.com/docs/en/goal) command gives you a generator/evaluator loop out of the box: set a completion condition and a separate fast model checks it after every turn until it's met. This repo ships the same underlying primitives as short, readable [hooks](https://code.claude.com/docs/en/hooks) and a [subagent](https://code.claude.com/docs/en/sub-agents), so you can see how each mechanism works and assemble a harness tuned to your project. The patterns come from [Effective Harnesses for Long-Running Agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) (Nov 2025) and [Harness Design for Long-Running Application Development](https://www.anthropic.com/engineering/harness-design-long-running-apps) (Mar 2026). We recommend trying both the in-product features and a custom harness to see which fits your workflow.\n\n| | In-product | Custom harness (this repo) |\n|---|---|---|\n| **What runs the loop** | [`/goal`](https://code.claude.com/docs/en/goal) | the [primitives below](#the-quality-loop) + a [loop you write](#running-the-loop) |\n| **Who judges \"done\"** | a separate fast model checking your condition | your [`agents/evaluator.md`](./claude-code-config/.claude/agents/evaluator.md) with your prompt |\n| **Where it works** | Claude Code interactive, [`-p`](https://code.claude.com/docs/en/headless), Remote Control | Claude Code, headless, or [Agent SDK](https://docs.claude.com/en/docs/claude-code/sdk) |\n\nThree primitives form the quality loop:\n\n- **Default-FAIL contract.** Every criterion starts `false`; the agent can't mark it passing without opening evidence first.\n- **Fresh-context evaluator.** A separate agent with no Write/Edit tools grades the work from a context window that never saw the build.\n- **Agent-maintained handoff.** The agent writes its own progress notes and commits to git so the next session picks up cleanly.\n\nTwo more operator-control hooks are included for when you want to watch or intervene. The same patterns translate directly to `PreToolUse`/`Stop` callbacks in the [Agent SDK](https://docs.claude.com/en/docs/claude-code/sdk).\n\n> Built as the take-home for the Long-Running Agents station at Code with Claude 2026. **These are example ingredients, not a turnkey harness.** Event demo; not maintained and not accepting contributions.\n\n## How to use this repo\n\n**Read and cherry-pick.** Each primitive is one standalone file with no dependency on the others. The [quality-loop table](#the-quality-loop) below maps every one to its example here and its Agent SDK equivalent. Open the file, see how the mechanism works, copy what fits.\n\n**Or copy all of them as a starting point.** In your project, run `claude` and paste:\n\n> Clone github.com/anthropics/cwc-long-running-agents into /tmp, copy its `claude-code-config/.claude/` directory into this project's root, make the hook scripts executable, then walk me through what each hook and the evaluator subagent does and what I'll need to adapt for this project.\n\nor do it yourself:\n\n```bash\ncp -r claude-code-config/.claude /path/to/your/project/\nchmod +x /path/to/your/project/.claude/hooks/*.sh\n```\n\nEither way, this gives you all the examples wired into `.claude/` at once. Before relying on them: point `RESULTS_FILE` at your project's actual results file, adjust the evidence-file pattern in `track-read.sh`, and run `claude` from the directory that contains `.claude/` (hooks are not loaded when launching from a subdirectory).\n\n**If you're on the Agent SDK,** this repo is a pattern reference. The shell hooks here translate one-to-one to `PreToolUse`/`Stop` callbacks. To scaffold an SDK agent from inside Claude Code, install the [`agent-sdk-dev` plugin](https://github.com/anthropics/claude-plugins-official/tree/main/plugins/agent-sdk-dev) and ask Claude to build an agent that implements whichever of these primitives you want. For a hand-written starting point, see the [autonomous-coding quickstart](https://github.com/anthropics/claude-quickstarts/tree/main/autonomous-coding) or the [agent-sdk-workshop](https://github.com/anthropics/agent-sdk-workshop) curriculum.\n\n## The quality loop\n\n| Primitive | Claude Code example | Agent SDK equivalent | Enforcement |\n|---|---|---|---|\n| **Default-FAIL contract** | [`hooks/track-read.sh`](./claude-code-config/.claude/hooks/track-read.sh) + [`hooks/verify-gate.sh`](./claude-code-config/.claude/hooks/verify-gate.sh) | `PreToolUse` callback | hook |\n| **Fresh-context evaluator** | [`agents/evaluator.md`](./claude-code-config/.claude/agents/evaluator.md) subagent (no Write/Edit) | [`evaluator_optimizer.ipynb`](https://github.com/anthropics/claude-cookbooks/blob/main/patterns/agents/evaluator_optimizer.ipynb) | you invoke it |\n| **Agent-maintained handoff** | [`CLAUDE.md`](./claude-code-config/.claude/CLAUDE.md) + [`hooks/commit-on-stop.sh`](./claude-code-config/.claude/hooks/commit-on-stop.sh) | system prompt + `Stop` callback | convention + hook backstop |\n\n### Default-FAIL contract\n\nAgents will mark a feature \"passing\" after a unit test or a curl when the UI is visibly broken. Asking nicely in the prompt doesn't reliably stop this. The harness makes \"done\" structural. Every feature is a row in a `test-results.json` file you create in your project:\n\n```json\n{ \"feature-1\": { \"passes\": false }, \"feature-2\": { \"passes\": false } }\n```\n\nThe only evidence that counts is a file matching the patterns in `track-read.sh` (screenshots, console logs, result files), and a `PreToolUse` hook denies any write to the results file unless the agent has first opened one with the Read tool. The agent can't claim success it hasn't observed. (The shipped hook is intentionally simple; see the comments in `verify-gate.sh` for the gaps a production version would close.)\n\n### Fresh-context evaluator\n\nThe builder shouldn't grade its own work. After each feature, you (or your wrapper script) invoke a separate subagent (`agents/evaluator.md`) with no Write/Edit tools that reviews the diff and the screenshots from a context window that never saw the build, then returns `PASS` or `NEEDS_WORK` with specific findings. On `NEEDS_WORK` the findings become the next builder session's starting prompt, closing the build/evaluate/rebuild loop. Invoke it from your wrapper with `claude --agent evaluator -p \"<review prompt>\"`; what your loop does with the verdict is up to you.\n\n### Agent-maintained handoff\n\nA fresh session has no memory of what the previous one did, and when a long session fills its context window Claude Code summarizes the history, which loses detail. So the agent maintains the handoff itself: it scopes each session to one feature, writes to a structured `PROGRESS.md` as it works and re-reads it first thing on every restart, and `git add`s and commits at meaningful checkpoints so `git log` is a second record. `commit-on-stop.sh` is the backstop that catches whatever's still uncommitted at session end. This is the layer most sensitive to model capability; newer models drift less and self-scope better, so re-evaluate how much of `CLAUDE.md` you still need after each model release (see [Re-simplify on model upgrades](#going-further)).\n\nA fourth core piece, a **rubric for subjective work**, isn't shipped here because it's project-specific; see [Going further](#going-further) for how to add one to the evaluator.\n\n## Running the loop\n\nTwo ways to keep the build → evaluate → rebuild cycle going. `/goal` is built into Claude Code and works with or without this repo's primitives; the second path wires the contract file and your own `evaluator.md` directly into the loop.\n\n### `/goal`: built-in completion checker\n\n```\n/goal every feature in PROGRESS.md is implemented, committed, and its tests pass\n```\n\nAfter every turn a separate fast model checks the condition and keeps the session going until it's met. One line, no contract file or hooks. Works the same in interactive Claude Code, [`claude -p`](https://code.claude.com/docs/en/headless), and Remote Control. See the docs for [writing an effective condition](https://code.claude.com/docs/en/goal#write-an-effective-condition) and how `/goal` [compares to `/loop` and Stop hooks](https://code.claude.com/docs/en/goal#compare-to-other-autonomous-workflows).\n\n### Your `evaluator.md` as the gate\n\nWhen you want the default-FAIL contract (`test-results.json` + `verify-gate`) enforcing evidence and your own evaluator prompt deciding `PASS`/`NEEDS_WORK`, the loop lives in your code. How depends on where you're running:\n\n| Surface | How |\n|---|---|\n| **Claude Code** | custom [Stop hook](https://code.claude.com/docs/en/hooks#stop) runs the evaluator as a fresh `claude` process after each turn and blocks on anything but `PASS` |\n| **Headless** ([`claude -p`](https://code.claude.com/docs/en/headless)) | wrapper script calls `claude --agent evaluator -p` between builds (example below) |\n| [**Agent SDK**](https://docs.claude.com/en/docs/claude-code/sdk) | separate generator and evaluator `query()` calls; see [`evaluator_optimizer`](https://github.com/anthropics/claude-cookbooks/blob/main/patterns/agents/evaluator_optimizer.ipynb) |\n\n```bash\nwhile grep -q '\"passes\": false' test-results.json; do\n  claude -p \"Read PROGRESS.md and build the next unfinished feature per CLAUDE.md.\"\n  VERDICT=$(claude --agent evaluator -p \"Review the most recent commit against its spec.\")\n  [ \"$(echo \"$VERDICT\" | head -1)\" = \"PASS\" ] || echo \"$VERDICT\" > NEXT_FINDINGS.md\ndone\n```\n\nEach pass is a fresh context. Exit when the contract file has nothing left failing, a cycle makes no changes, or a budget is hit; `touch AGENT_STOP` to stop early. The builder writes `test-results.json` (verify-gate enforces evidence-read first) and the evaluator returns a separate verdict; have your wrapper write the contract on `PASS` instead if you want \"all true\" to mean \"independently confirmed.\"\n\n## Operator controls\n\nTwo more hooks for when you want to watch the loop run or intervene mid-stream:\n\n| Hook | What it does |\n|---|---|\n| [`hooks/kill-switch.sh`](./claude-code-config/.claude/hooks/kill-switch.sh) | Halts every tool call while an `AGENT_STOP` file exists at the project root. |\n| [`hooks/steer.sh`](./claude-code-config/.claude/hooks/steer.sh) | Surfaces the contents of `STEER.md` (project root) to the agent once and clears it, so you can redirect mid-run without restarting. |\n\n## Watching it work\n\nEverything the agent does lands on disk, so you can observe a long run from a terminal without any dashboard code:\n\n```bash\nwatch -n 2 'tail -20 PROGRESS.md'                          # its own notes\nwatch -n 5 'git log --oneline -8'                          # work saved\nwatch -n 5 'find screenshots -name \"*.png\" | tail -5'      # what it sees\nwatch -n 2 'wc -l < .claude/.evidence-reads 2>/dev/null'   # evidence reads (resets to 0 after each gated write)\n```\n\nRun them in a `tmux` grid and you have a live view of every primitive.\n\n## Going further\n\nThe next layer of patterns, each covered in depth in [Harness Design for Long-Running Application Development](https://www.anthropic.com/engineering/harness-design-long-running-apps):\n\n| Pattern | What it adds | Reference |\n|---|---|---|\n| **Unattended loop** | Cap session length and have an outer script start the next one (pick next feature, build, evaluate, reset) | [Why naive implementations fall short](https://www.anthropic.com/engineering/harness-design-long-running-apps#why-naive-implementations-fall-short); [`ralph-loop`](https://github.com/anthropics/claude-plugins-official/tree/main/plugins/ralph-loop) plugin |\n| **Planner agent** | A first session that expands a one-line ask into a `BUILD_PLAN.md` the loop then runs against | [The architecture](https://www.anthropic.com/engineering/harness-design-long-running-apps#the-architecture) |\n| **Sprint contracts** | Builder and evaluator agree per-feature on what \"done\" means and write it to a file the hook enforces | [The architecture](https://www.anthropic.com/engineering/harness-design-long-running-apps#the-architecture) and [Removing the sprint construct](https://www.anthropic.com/engineering/harness-design-long-running-apps#removing-the-sprint-construct) |\n| **Grading rubrics** | Give the evaluator scoring principles (functionality, design, craft, originality) with few-shot examples instead of binary pass/fail | [Frontend design: making subjective quality gradable](https://www.anthropic.com/engineering/harness-design-long-running-apps#frontend-design-making-subjective-quality-gradable); [`frontend-design`](https://github.com/anthropics/claude-plugins-official/tree/main/plugins/frontend-design) plugin |\n| **Browser-verified evaluator** | Let the evaluator open the running app itself instead of trusting the builder's screenshots | [Frontend design](https://www.anthropic.com/engineering/harness-design-long-running-apps#frontend-design-making-subjective-quality-gradable) (Playwright MCP usage); add [`@playwright/mcp`](https://github.com/microsoft/playwright-mcp) or Claude in Chrome to `tools:` in `agents/evaluator.md` |\n| **Re-simplify on model upgrades** | After each model release, comment out harness pieces one at a time and see what's still load-bearing | [What comes next](https://www.anthropic.com/engineering/harness-design-long-running-apps#what-comes-next); [Harnessing Claude's Intelligence](https://claude.com/blog/harnessing-claudes-intelligence) |\n| **Hosted runtime** | Anthropic hosts the loop, sandbox, and scheduling so you don't run any of this yourself | [Claude Managed Agents](https://docs.claude.com/en/docs/managed-agents) |\n\n---","is_error":false}]},"uuid":"13741e26-ae5e-41dd-95b0-e52335c1518f","timestamp":"2026-07-09T08:00:44.421Z","sourceToolAssistantUUID":"2e967a83-f54c-4126-b259-0f26f56668e1","userType":"external","entrypoint":"sdk-ts","cwd":"/workspace/agent","sessionId":"033a4a00-9171-4e08-9a79-6b99f177355d","version":"2.1.128","gitBranch":"HEAD"}
{"parentUuid":"13741e26-ae5e-41dd-95b0-e52335c1518f","isSidechain":true,"agentId":"ab62bea074e7ddd95","message":{"model":"claude-haiku-4-5-20251001","id":"msg_011Ccr8bUgoCSB5BMYB6P4fe","type":"message","role":"assistant","content":[{"type":"text","text":"## Summary\n\n**1. Claude SDK for Go** — Official Go language SDK for the Anthropic Claude API\n- Provides full access to the Claude API from Go applications\n- Install: `go get -u 'github.com/anthropics/anthropic-sdk-go@v1.56.0'`\n- Creates client with API key, then calls `client.Messages.New()` for chat interactions\n- Simple, type-safe interface; requires Go 1.24+\n- Stands out: Clean abstractions for message params and model selection (e.g., `ModelClaudeOpus4_6`)\n\n**2. CWC Long-Running Agents** — Educational workshop repo shipping production patterns for autonomous Claude agents\n- Three core primitives for building reliable long-running agentic loops: default-FAIL contract (evidence-based verification), fresh-context evaluator (separate agent grades work), agent-maintained handoff (agent writes progress, commits to git)\n- Use patterns as standalone files (cherry-pick what fits) or copy entire `.claude/` hook setup; works with Claude Code interactive, headless (`claude -p`), or Agent SDK\n- Key feature: `/goal` built-in completion checker runs evaluator automatically; or wire your own `evaluator.md` subagent + `test-results.json` contract file\n- Stands out: Not a turnkey tool but educational reference from Anthropic's long-running agents research; teaches structural enforcement of \"done\" (can't claim success without opening evidence first)"}],"stop_reason":"end_turn","stop_sequence":null,"stop_details":null,"usage":{"input_tokens":7,"cache_creation_input_tokens":6487,"cache_read_input_tokens":8438,"output_tokens":341,"server_tool_use":{"web_search_requests":0,"web_fetch_requests":0},"service_tier":"standard","cache_creation":{"ephemeral_1h_input_tokens":0,"ephemeral_5m_input_tokens":6487},"inference_geo":"","iterations":[{"input_tokens":7,"output_tokens":341,"cache_read_input_tokens":8438,"cache_creation_input_tokens":6487,"cache_creation":{"ephemeral_5m_input_tokens":6487,"ephemeral_1h_input_tokens":0},"type":"message"}],"speed":"standard"},"diagnostics":null},"requestId":"req_011Ccr8bToDWZ2yZyiq2oih2","attributionAgent":"Explore","type":"assistant","uuid":"136ddd31-11e3-4f05-8b01-4d285275a9a0","timestamp":"2026-07-09T08:00:51.060Z","userType":"external","entrypoint":"sdk-ts","cwd":"/workspace/agent","sessionId":"033a4a00-9171-4e08-9a79-6b99f177355d","version":"2.1.128","gitBranch":"HEAD"}
