OpenAI's GPT-6 Astra can now do tax returns and CAD work — enterprises should read the fine print before deploying it
OpenAI released GPT-6 Astra, which president Greg Brockman called a "generational leap" and said may mark the arrival of AGI. The model works directly inside software — drafting tax returns from a W-2, laying out circuit boards, building 3D scenes — rather than just giving advice.
Astra is the first OpenAI model to hit the company's "critical" cybersecurity threshold under its preparedness framework, meaning it can find and exploit unknown vulnerabilities without step-by-step guidance. It's also harder for OpenAI's own researchers to monitor than earlier models — a decline the company calls serious even as it insists the model still can't fully hide its reasoning. For enterprises moving agents from pilot to production, capability and control are now separate questions, and the safety-testing lag between announcement and general availability is where the real work happens.
An agent that files a draft tax return or lays out a PCB is doing real professional work with real liability if it's wrong — verify outputs, don't just skim them.
Engineers: Astra's harder-to-monitor reasoning means your evals and guardrails need to assume less visibility into "why," not just "what" — chain-of-thought monitoring alone won't cut it.
Finance: A model that can draft a tax return from a W-2 is a preview of where compliance and audit risk is heading — treat AI-drafted filings as a first draft requiring sign-off, not a submission.
Managers: Rolling this out to teams means defining what "autonomous" actually covers — set explicit boundaries on what Astra can execute versus merely propose.
Do this: Before letting Astra (or any agent at this capability tier) touch production tax, legal, or engineering work, require a human sign-off step and ask your vendor what monitorability testing they've done.