Nobody's Talking About the Real Story in Devin and Perplexity's AI Upgrades

Creative Robotics
Nobody's Talking About the Real Story in Devin and Perplexity's AI Upgrades

Two announcements this week should have set off alarm bells, but instead they arrived with barely a whisper of concern. Perplexity deployed GPT-6 Astra for "end-to-end systems" management—writing communications, modifying software, monitoring production. Cognition integrated the same model into Devin so the AI coding assistant could test its own work. Read those sentences again. We're not talking about AI helping humans anymore. We're talking about AI validating AI.

The framing in both announcements focuses on efficiency gains and reduced workload. Perplexity needs less human oversight. Cognition's engineers spend less time on code review. Faster shipping cycles. Higher productivity. All true, and all beside the point.

What's actually happening is a fundamental shift in how software gets built and deployed. For decades, the iron law of development has been human review. You write code, someone else checks it. You modify a production system, another engineer signs off. You ship a feature, QA validates it works. These weren't just best practices—they were guardrails against catastrophic failure.

Now we're removing them, and replacing them with... what, exactly? More AI. The same technology that introduced the errors is now trusted to catch them. It's circular in a way that should make anyone uneasy, yet the industry is sprinting toward it anyway.

The optimistic view holds that GPT-6 Astra represents such a leap in capability that it's fundamentally different from earlier models. Maybe it is. The Cognition announcement specifically touts "enhanced automated testing" and better code validation. Perplexity clearly believes the model is reliable enough to modify production systems unsupervised. These aren't small companies making reckless bets—they're sophisticated engineering organizations.

But capability isn't the whole story. When AI tests AI-generated code, we create a closed feedback loop with no external validation. Errors that both the generator and validator share become invisible. Systematic biases get reinforced rather than caught. Edge cases that neither model anticipates simply don't exist in the testing regime.

Worse, we're normalizing a development culture where "it passed the AI review" becomes sufficient justification for deployment. Junior engineers won't learn to spot subtle bugs because the AI already approved the code. Senior engineers won't catch systematic issues because they're reviewing AI summaries rather than actual changes. The institutional knowledge of what good code looks like—and more importantly, what bad code looks like—gradually erodes.

This isn't hypothetical. We've seen this movie before with automated testing suites that gave false confidence, with CI/CD pipelines that shipped broken builds faster than manual processes ever could, with monitoring systems that missed the exact failure modes they weren't programmed to detect.

The difference is scale and speed. When Perplexity's AI modifies production systems autonomously, failures propagate at machine speed. When Devin tests its own work and ships based on that self-assessment, entire codebases evolve without meaningful human comprehension of what changed or why.

Maybe GPT-6 Astra really is good enough to make this work. Maybe the efficiency gains are worth the risks. But let's at least be honest about what we're doing: we're building a software industry where the machines increasingly check each other's work, and we're calling it progress because it's faster. Whether it's actually better remains an open question we seem determined not to ask.