Prompts for testing what new AI models can actually build.

A small, opinionated collection of personal benchmarks. Same brief, different model, visible result.