Background
The trigger was several one-on-one meetings with my team lead. He was interested in how I use AI tools, build personal Harnesses, and structure my AI Coding workflow. Mainly because I am delivering requirements quickly and well, already serving as the primary owner for some projects. So, I took this opportunity to organize my daily AI workflow. Currently, I mainly do Agent development in my group while maintaining existing backend business logic.
Previously, I shared related workflows on Xiaohongshu. The context then was that GPT-5.4 and Opus 4.6 could already complete tasks well when given accurate context and reasonable constraints. The rise of Harness Engineering made everyone realize: you can better control powerful models by adding constraints.
For example, Skills like Superpowers provide a relatively heavy implementation, mainly revolving around Spec and TDD, helping people organize and advance tasks more easily.
But doing so has some drawbacks. The most obvious feeling is: Token consumption is very fast. Skills like Superpowers contain many Workflows used to constrain the model's next action. Even if, from a global perspective, a certain step is no longer necessary—for instance, the context is sufficient and implementation can start directly—the model may still continue executing along the predetermined process.
With the release of new models, I have seen many developers start sharing that they basically no longer use these very heavy Skills. One main reason is that after model capabilities enhance, some constraints and processes we originally thought were useful are becoming noise for the model. For example, when OpenAI released Astra, they specifically wrote a blog post introducing how to clean up unnecessary Skills and system prompts to improve the model usage experience.
So, this article will combine my practices over the past few months to share some methods that are currently effective for me, discussing how to better use various Agents to improve daily development efficiency.
1. My Commonly Used Skills and Prompts
My current tool division:
Currently, I mainly use Codex + GPT-5.6 Sol for coding implementation, and Astra for planning. Before Astra was released, planning work was mainly handed to GPT-5.6 Sol Max.
For simple requirements, I use pi + DeepSeek V4 Flash for implementation; adversarial review of solutions and Code Review are mainly handled by Claude 5 Fable. In personal projects, I also use the web-based GPT-6 Pro for main solution design.
My Commonly Used Skills
- think: tw93's Skill, mainly used for solution alignment and brainstorming.
- grill me / grill with docs: Mainly used for requirement clarification. Through continuous questioning, it clarifies goals, constraints, and trade-offs, then precipitates them into ADR or CONTEXT.md according to development process needs. It helps me dig out issues not considered during early alignment.
- implement: Used in conjunction with grill, part of Matt Pocock's Skills suite, used to implement clearly defined plans or Issues.
- ponytail: Used to clean up AI over-design, reducing Review difficulty. I use this often because GPT-5.6 Sol frequently over-designs.
- handoff: Organizes current context into files, facilitating task continuation in new sessions. I generally use it to hand off from Codex to Claude Code or pi agent.
- check: Used for code review, generally used when submitting MRs.
- Skills precipitated from work and personal development: Mainly reusable process SOPs, such as end-to-end testing. My suggestion is that in daily work, if a process repeats more than three times, consider having Codex help organize it into a Skill for direct reuse later.
My Commonly Used Prompts
I rarely write large chunks of Prompt myself now. When needed, I usually let Codex help organize them. For example, after multiple rounds of discussion, if I want to hand the current context to GPT Pro for solution design, I have Codex generate a complete handoff Prompt first.
Besides this, I frequently use the following types of very short Prompts.
They sometimes achieve the effect of "one sentence worth thousands." I understand them as the model's "thinking shortcuts."
They are not magic spells, but methodologies highly standardized in human knowledge. During training, the model has seen a large amount of related papers, code, design documents, and discussions, so often there is no need to hand-write hundreds of lines of Workflow, just tell it which thinking method to adopt. Here are some prompts I found very effective through practice:

- First Principles: Do not continue optimizing along existing solutions, re-ask how this problem should actually be solved. For example, if an interface is slow, you can say:
Do not continue designing based on the established solution of "adding Redis cache." Analyze from first principles why this interface is slow, and what the minimum necessary solution is.
The model's focus shifts from "how should Redis be designed" to:
Is the bottleneck in SQL, network, serialization, lock contention, or repeated calculation? If adding an index to SQL solves it, why introduce Redis?
This type of Prompt is suitable when you suspect "the question itself might be wrong."
- Adversarial Review: Do not find reasons for my solution, try to prove it is wrong.
The ordinary way to ask is:
Help me see if there are any problems with this technical solution.
It can be changed to:
Conduct an adversarial review of this solution, prioritizing finding counter-examples that can overturn core assumptions.
Assuming your solution is "introducing distributed locks to solve duplicate requests," the model will no longer just tell you how to set the lock timeout, but start asking:
Do duplicate requests really need mutual exclusion? Can interface idempotency solve it? What if the lock service crashes? What if the lock expires but the business hasn't finished executing? Did we introduce a new distributed failure point to solve a local problem?
This type of Prompt is particularly suitable for solution reviews and Code Reviews.
- Ablation Experiments: The system getting better does not mean every thing you added is useful.
For example, you made three optimizations at once:
After adding indexes, Redis cache, and batch queries, interface latency dropped from 800ms to 100ms.
At this time, you can directly ask:
Design ablation experiments for these three optimizations to determine where the real benefits come from.
The model will design controls around different combinations, starting from Baseline, comparing only adding indexes, indexes + cache, indexes + cache + batch queries, etc.
Finally, it might find:
Only adding indexes already reduced latency from 800ms to 120ms, the remaining two complex solutions contributed only 20ms.
This way, you know more clearly which code is worth keeping and which complexity might be unnecessary.
- Occam's Razor: When effects are similar, prioritize solutions with fewer assumptions and lower complexity.
For example, the Agent designs a solution like this:
Kafka + Redis + Distributed Lock + State Machine + Timed Compensation.
You can add a sentence:
Re-review this design using Occam's Razor, deleting all non-essential mechanisms under the premise of meeting requirements.
Often, it finally finds:
The current scenario only involves single-database writes, one transaction plus a unique index is enough.
This sentence is especially useful for current Coding Agents, because models easily over-design for the sake of "completeness."
- High Cohesion, Low Coupling: Re-check code responsibilities and boundaries.
For example, if you find an OrderService already has 2000 lines, you can ask:
Review the responsibility boundaries of OrderService according to high cohesion and low coupling principles, do not split for the sake of splitting.
The model usually starts checking:
Why does the order service simultaneously handle inventory, coupons, SMS, payment, and reports? Which logic belongs to the order domain itself, and which should be handed to other modules via stable interfaces?
It activates not just simple "file splitting," but a whole set of judgments about modularization, information hiding, dependency direction, and responsibility division.
So, I rarely write anymore:
Step 1 analyze requirements, Step 2 check assumptions, Step 3 find alternatives, Step 4...
Many mature methodologies have already been learned by the model. I prefer to directly tell it:
Re-analyze from first principles, conduct adversarial review on the current solution; core mechanisms must be verified through ablation experiments; the solution follows Occam's Razor, code maintains high cohesion and low coupling.
Behind these dozens of words, five different cognitive actions are actually specified:
Redefine the problem → Attack assumptions → Verify contributions → Delete complexity → Organize system boundaries.
This is also a change in Prompt Engineering in the new model era as I understand it: Instead of writing a complete fixed thinking process for the model, it is better to use accurate methodologies to tell it "what way to think," and then supplement the truly necessary constraints for the current task.
2. My Daily Development Workflow

After receiving a requirement, I usually first hand relevant context to the Agent, such as PRD, meeting minutes, chat records, and user feedback, then use grill to align requirements with it.
These materials are often not a complete, consistent requirement. The PRD may not be updated, some restrictions may have been supplemented in meetings, priorities adjusted in chats. My own understanding of requirements may also include some unstated assumptions.
I let the Agent understand these materials combined with team Wiki and existing code, then clarify goals, boundaries, and trade-offs affecting implementation through continuous questioning. Some questions I can answer on the spot, others require going back to confirm with product or relevant colleagues.
Here, I control a scale: when the remaining questions will no longer significantly change the implementation direction and acceptance results, development can begin.
I do not require it to plan all implementation details in advance, otherwise requirement alignment itself becomes a very heavy process.
Conclusions after alignment are precipitated into Spec or CONTEXT.md, mainly recording the problem to be solved this time, scope, key decisions, and acceptance criteria. This also facilitates subsequent Agents responsible for implementation and review to share the same context, without needing to re-read all previous conversations.
After the solution is determined, I let the main Agent decide how to execute based on task complexity. Simple requirements are implemented directly; complex requirements are split into Issues with clear boundaries that can be independently accepted. Only parts that can proceed independently are handed to subagents for parallel development in different worktrees, finally integrated by the main Agent.
The coding Agent completes tests first. When it thinks it is ready to deliver, I introduce other Agents for cross-adversarial review based on task complexity. Discovered issues are uniformly fed back to the main coding Agent, i.e., Codex, which fixes and re-verifies.
Before my own Review, I also perform end-to-end testing.
I currently do not read all generated code line by line, mainly looking at test results and core business logic. There is an important prerequisite here: Each MR scope is small enough, and complete business paths are gradually verified as development progresses.
Small MRs keep the changes requiring understanding and judgment within controllable ranges each time; end-to-end tests help check whether these changes hold up when placed into real business processes. During manual Review, I focus on confirming business logic and whether existing test results are sufficient to support this delivery.
3. How to Build Agent-Friendly End-to-End Testing
AI is already very good at writing test cases, often writing hundreds of lines of tests for a Bugfix (especially 5.6 sol). Often, writing more tests itself is not a problem, but after deployment and launch, we still encounter unexpected errors.
In my practice, a core issue is: Not providing the Agent with an end-to-end testing environment, giving it the opportunity to discover these problems during coding and self-testing stages.
If such an environment can be built, allowing the Agent to conveniently execute tasks from the real entry point of the test environment, checking all the way to the results users need, we can confidently let AI-written code run in real systems.
In actual development, I have built a set of end-to-end development testing environments for our business. The Agent can conveniently query database tables, logs, and connect to machines to troubleshoot issues.
The entire building process is essentially letting the Agent extract my daily development self-test processes, integrating scattered tools and capabilities: Capabilities that can be toolized are made into MCP or CLI; reusable processes are written into Skills.
This indeed requires some effort upfront, but do not fear trouble. Once built, it significantly speeds up development, reduces rework and online issue possibilities, and you don't have to worry all day about whether AI-written code will cause accidents.

Around Agent-friendly end-to-end testing, I mainly did four things:
- Let the Agent familiarize with the environment: Such as supporting one-click startup of test environments, creating test data, clarifying current versions, test accounts, and permissions, and providing cleanup and reset capabilities.
- Let the Agent operate business systems: Execute real business processes via browsers, APIs, or CLIs.
- Let the Agent conveniently query database tables and logs: Confirm storage results via read-only database MCP, locate issues via log system Skills and Trace queries.
- Reuse tedious but stable processes: Write stable operations into scripts, organize entry points and troubleshooting methods into Skills, reducing repetitive manual intervention and dialogue.
Based on these works, here are a few methods I currently find effective:
- Browser Automation: When business systems require browser operations, I recommend the open-source browser ego lite. It is convenient to use and fast. Paired with pi agent + DeepSeek V4 Flash, tests can be completed relatively quickly, reducing test time.
- Unify Operations into CLI: Integrate reusable Skills, configured MCPs, and written scripts into a unified test CLI, providing capabilities for environment checks, data preparation, scenario execution, result queries, and cleanup. This can also become an internal efficiency tool. Currently, I have made this set into a CLI, and it is also convenient for troubleshooting when encountering issues.
- Let Tools Directly Serve Acceptance: DB queries are used to confirm status, logs and Traces are used to explain failures, but expected results must still come from business contracts. Do not let the Agent assume what is correct just because it sees what the system returned. Under complex business logic, the system may completely return a seemingly reasonable but actually non-compliant result. Tools help us obtain evidence, not define correct answers for us.
- If the Product Itself is an Agent, Also Check Answer Quality: If the product being tested is itself an Agent, besides whether business processes run through, answer quality also needs consideration. This is what we often call Agent Eval, which won't be expanded here.
4. How to Review AI-Written Code
Previous sections shared how to make AI write high-quality code, but ultimately, the developer remains the primary person responsible for business requirements.
Without Review, large systems easily encounter problems.
You also don't want to be called up for On-call in the middle of the night, only to find out it was AI-written code causing issues.
Regarding Review, I currently have the following main practices:
- Look at acceptance criteria first, then test results: First check the acceptance criteria given by the Agent, then compare with previous end-to-end test results to confirm if expectations are met and if anything is missed. Don't just look at how many tests passed, but also see if these tests verified what this requirement truly cares about.
- Follow business paths to view implementation, focusing energy on high-risk parts: I mainly check permissions, state changes, concurrency, retries, data consistency, and high-risk parts like migration and rollback. For stable pattern CRUD, less time can be spent—part of this I even don't look at now.
- Specifically check newly added abstractions and mechanisms: For newly added abstractions and mechanisms, I use thoughts like Occam's Razor to let AI review again: Are they truly necessary, is there a simpler implementation, did they introduce too much complexity for local problems?
- Consider Building Dedicated Review Bots When MRs Are Numerous: If there are many MRs in the group, a dedicated Review Bot can be designed to handle cross-review. Different from directly calling pi agent / Claude Code for adversarial review mentioned earlier, it emphasizes combining Git change information with pre-designed review processes, forming a repeatable review capability specifically serving MRs.
5. Some Summaries and Thoughts
My most obvious bottleneck in current development is Review speed.
Agents can push forward several tasks simultaneously, but my speed of understanding business, judging solutions, and confirming deliveries does not increase proportionally. If you just let it write more, you likely just accumulate more code waiting for review.
So, what I want to improve next is letting repetitive issues be discovered and fixed before reaching my hands.
Issues that can be found through type checking, tests, and business assertions should be handled by the Agent itself during development as much as possible; those requiring my judgment are concentrated on whether requirements are understood correctly, whether key business logic holds, and what unverified risks remain in this change.
Dedicated Review Bots can help do these things, but their value depends on whether they reduce omissions of effective issues and manual burden, not just how many comments they make.
This also makes my understanding of Harness gradually concrete: Besides giving Agents correct context, they also need environments capable of executing tasks, and bases for judging results.
I organized operations repeatedly done in daily self-tests into CLIs, scripts, and Skills, letting it run systems, check results, and find failure evidence itself. In future similar tasks, these capabilities can continue to be used, and gradually handed to other colleagues for reuse.
Meanwhile, this workflow also needs regular subtraction.
Some steps are compensating for shortcomings of a certain generation of models. When models change, these steps' benefits should be re-evaluated. Simple tasks are done directly, complex tasks add planning, splitting, and cross-review. For example, after Astra updates, I deleted some overly strict constraints in AGENTS.md. Models iterate, and our workflows need to iterate accordingly.
Of course, testing and Review can reduce uncertainty, but humans still need to give correct acceptance criteria. Even if code and tests match each other, they might both misunderstand the requirements together. End-to-end tests only cover behaviors in selected environments and scenarios; traffic, concurrency, and data distribution in production environments may still bring new problems.
For a developer who has just started working like me, I still hope to learn more professional knowledge in the field through work. However, AI has indeed reduced some opportunities to personally step into pitfalls. Some valuable experiences originally gained through experiencing problems and investigating causes may now turn into an AI saying:
"I was wrong, modifying now."
So,
I now reserve part of daily development time for learning and reflection, while thinking: What abilities are truly needed for R&D colleagues in the AI era.
This article records a set of methods I gradually explored in my business during my first few months. Its scope of application and deficiencies, I am still exploring.
Everyone can first pick a self-test path commonly done, try letting the Agent run it independently, preserve evidence, then precipitate effective steps.
Due to company confidentiality requirements, many details cannot be included in the article. Hope this article throws a brick to attract jade, and I'd love to hear good practices from everyone in actual development~





