I've been using GPT-6 Astra with the workflow I developed around GPT-5.6: documented task scope, repository context, standard operating procedures, and verification before delivery. In my testing, I've had to keep correcting work those instructions already covered. Examples included pointing back to a supplied writing guide, reasserting testing cadence, and asking for requested work missing from a completion handoff.
Read the
full study , the
blog post , or explore the
data and methods in the research repo.

25,254 recorded responses across six study days. Day 6 ends at the common cutoff. Source: Astra Field Study usage records.
I reviewed six days of retained Astra records to understand the activity and the steering around it. The records contained 25,254 responses across 557 turns. A separate review of my typed and voice contributions identified 147 corrective contributions among 538 substantive contributions, grouped into 78 episodes. Thirty-one episodes contained repeated corrections.
The process predates Astra
I wrote about ongoing Codex work, task scope and constraints, and versioned SOPs in June. Eric Provencher from OpenAI's Codex DX team described closely aligned practices in his July guidance for GPT-5.6, including distinct assignments and preserving requirements when delegating.
My established method hasn't carried over reliably in my Astra testing. The instructions cover access, testing, and delivery requirements for the work.
What the review shows
Forty correction episodes involved process or SOP concerns. One testing-related episode contained 17 corrective contributions. The review also recorded 154 contributions expressing dissatisfaction, with 91 overlapping the corrective category. The same problem sometimes generated several corrections, so this count doesn't measure a model failure rate.

538 substantive contributions in five exclusive categories: 310 routine, 63 dissatisfaction only, 56 correction only, 91 both, and 18 ambiguous. The 147 corrective contributions include the 91 in both categories. Source: contextual review, protocol 1.1.
The usage total was 3.42 billion recorded tokens, including 3.35 billion cached input tokens. About 98.2% of input came from cache. Repeated context contributes to those counters, so the total doesn't tell us how much useful work was completed or what it cost. The study includes the full breakdown and keeps repository activity separate.

Recorded token composition: 3.35B cached input, 61.11M uncached input, and 10.99M output tokens. Cached input accounts for 98.21% of input. Output includes reasoning; these counters do not measure cost or useful work. Source: Astra Field Study usage records.
Two hypotheses to test
I suspect shortcuts toward finishing can displace requirements, and requirements can get lost during delegation. Both are entirely untested hypotheses based on roughly four or five days of personal experience. I haven't run a matched GPT-5.6/GPT-6 comparison or traced a causal delegation failure. A deployment runbook contained outdated instructions.
The replies have included different experiences. Arbaz reported drift under detailed design instructions, while John Collins reported good instruction-following when steps were explained.
I've published Astra Field Study with reviewed aggregates, methods, figures, an explorer, and an MIT contribution kit. The kit keeps raw conversations private and lets contributors review their aggregate before submitting it. Optional satisfaction and ease ratings are collected directly and kept separate from usage.
This is what I currently think, and I want to know more. If your existing workflow transferred cleanly, share what you used and how you verified the result. Contribute through the repo or reach out here, especially if your experience contradicts mine.





