Ask a model whether it should be writing your production code
Here's an experiment that takes about two minutes. Ask a frontier model, plainly and without leading it, whether it should be writing production code with minimal human oversight in a system where failure carries real consequence. What comes back is a fairly consistent no. Some cases are fine: