You open Claude Code, describe a feature in one vague paragraph, and ask it to build everything.
The result is often a large diff that mixes discovery, architecture, implementation, cleanup, and testing into one difficult review.
You then correct one detail, receive another large diff, and slowly lose track of which decisions are still intentional.
The model may be capable, but the workflow gives it no stable contract and gives you no clean checkpoint.
Here is the practical rule: stop treating Claude Code as a chatbot that should solve the whole task in one answer.
Treat it as a small engineering team whose members need narrow roles, explicit handoffs, and gates before code reaches a pull request.
The giant-prompt workflow hides five different jobs
A normal feature request contains at least five jobs: define the outcome, implement it, review the diff, prove the acceptance criteria, and prepare the change for release.
When one conversation owns all five jobs at once, the same context that produced the code is also asked to judge that code.
That is convenient, but it weakens the separation between author and reviewer.
It also encourages scope drift because an implementation detail can quietly become a product decision without returning to you for approval.
A better workflow makes each job visible and forces the output of one stage to become the input of the next.
idea
↓
/spec
↓
/build
↓
/review
↓
/qa
↓
/ship
The important part is not the command names because the important part is that every stage has one owner and one definition of done.
Make the specification the contract
The specification should be short enough to review before implementation and precise enough to expose disagreements while they are still cheap.
A useful build spec names the requested behavior, acceptance criteria, constraints, likely files or components, validation steps, and deliberate exclusions.
The exclusions matter because they stop a small feature from turning into an unplanned architecture rewrite.
For example, a mobile feature request could begin like this:
/spec Add debounced repository search to the sample SwiftUI app.
Constraints:
- Keep the existing MVVM structure.
- Support iOS 13 and use Combine.
- Add no third-party dependency.
- Preserve the current empty and error states.
- Add unit tests for empty, loading, success, and failure results.
Out of scope:
- Search history.
- Offline caching.
- UI redesign.
This is still a compact request, but it gives the agent boundaries that can be reviewed before a file changes.
Once you approve the spec, the implementer can work against a stable document instead of repeatedly interpreting a moving conversation.
Give every stage one job
1. Spec writer: turn intent into acceptance criteria
The spec writer should inspect enough context to produce a realistic plan without starting the implementation.
Its output is a contract for the remaining stages and a checkpoint where you can correct direction with almost no rework.
2. Implementer: execute the approved scope
The implementer should make the smallest coherent change that satisfies the approved criteria.
When it discovers a requirement conflict, it should report the conflict instead of silently expanding the feature.
3. Reviewer: attack the diff, not the author
The reviewer should compare the implementation with both the spec and the surrounding codebase.
It should look for missed acceptance criteria, regressions, unsafe assumptions, unnecessary complexity, weak error handling, and changes outside the approved scope.
A separate reviewer role does not guarantee correctness, but it creates a fresh pass with a different objective.
4. QA: prove behavior against the contract
The QA stage should translate acceptance criteria into executable checks whenever the repository supports them.
Its final report should distinguish passing evidence, failing evidence, and anything that could not be verified.
A green statement without a command, test result, or reproducible check is confidence rather than evidence.
5. Ship: run a pre-flight check
The final stage should confirm tests, linting, documentation or changelog needs, commit state, and pull-request readiness.
This stage is a checklist rather than permission to merge blindly.
Use memory for conventions and hooks for enforcement
Claude Code supports project instructions, specialized subagents, reusable commands or skills, and lifecycle hooks, but those mechanisms solve different problems.
Use CLAUDE.md for architecture notes, build commands, naming conventions, and guidance that should load with the project.
Use deterministic hooks for rules that should run before or after a tool action, such as formatting an edited file or rejecting a known destructive shell command.
Anthropic's Claude Code feature overview makes the same practical distinction between instructions that guide behavior and hooks that enforce repeatable actions.
A guard hook is still defense in depth rather than a complete security boundary.
You should continue using repository permissions, protected branches, secret scanning, CI, backups, and human review according to the risk of the project.
This matters even more when a repository touches authentication, payments, personal data, production infrastructure, or destructive database operations.
What Ship It Solo packages for you
You can build this system yourself from the official Claude Code documentation and a collection of free examples.
The expensive part is deciding how the pieces hand work to each other, testing those decisions together, documenting the failure modes, and resisting a folder full of disconnected agent prompts.
Ship It Solo packages that wiring into 30 files built around one spec-to-PR pipeline.
- Four subagents: spec writer, implementer, reviewer, and QA.
- Five workflow commands:
/spec,/build,/review,/qa, and/ship. - Four guard and automation hooks: destructive-command blocking, obvious-secret blocking, automatic formatting, and a session brief.
- Five practical playbooks: new feature, bug hunt, refactor, test backfill, and release.
- A layered memory system: user, project, and feature templates plus an annotated example.
- Installation and failure documentation: an idempotent installer with dry-run support, manual installation steps, backups, and a guide to known failure modes.
The kit is stack-agnostic, while the included Bash scripts target macOS and Linux directly and Windows through WSL.
It works with a Claude Code subscription or API billing, but it does not work without Claude Code.
The model cost cascade assigns less expensive models to mechanical stages and stronger models to design or review, while the guide asks you to measure the result with your own /cost data.
Who should use it and who should skip it
This kit fits solo developers and indie hackers who already use Claude Code, ship real repositories, and want a repeatable workflow without designing every agent file and hook from scratch.
It also fits developers who have read about subagents and hooks but have never connected them into one reviewable process.
You should skip it if you want a no-code product, do not use Claude Code, or already have a mature agent pipeline with established CI and review conventions.
You should also skip it if you expect an agent reviewer to replace your responsibility for the final diff.
The real upgrade is the handoff
The biggest improvement does not come from adding more agents because it comes from giving each agent a smaller decision surface.
A reviewed spec narrows implementation, a narrow implementation creates a readable diff, a readable diff improves review, and explicit acceptance criteria give QA something concrete to prove.
You remain responsible for the architecture, the risk boundaries, and the final decision to ship.
The agents make that responsibility easier to exercise because they turn an open-ended chat into visible engineering stages.
Technical reference: Claude Code subagents and Claude Code hooks.
Comments