Researchers at Cornell have developed a new AI approach for generating more realistic images involving multiple people interacting with one another. While existing image-generation models can create convincing individual people, they often struggle to represent complex interactions accurately. The Cornell method addresses this through iterative pose-image generation, progressively constructing a scene one person at a time. Each predicted pose helps guide the generation of subsequent people, allowing the system to better capture spatial relationships and interactions without requiring users to manually specify poses. The approach uses the FLUX image-generation model as its foundation and combines pose detection with a multimodal large language model to organise descriptions, poses and spatial regions for each person.
The researchers also introduced DrawWaldoWorlds, a benchmark designed to evaluate whether image-generation systems correctly represent not only multiple individuals but also their roles and relationships, essentially testing whether the model understands who does what to whom. Experiments showed that the new approach produced more faithful multi-person scenes than existing methods. In a user study involving 20 participants, images generated using the Cornell method were preferred roughly two-to-one over images produced by two versions of FLUX. The researchers argue that automatically incorporating pose information could make generative AI considerably better at depicting complex social activities, sports, group scenes and other situations where realistic human interaction is essential.
More information:
https://news.cornell.edu/stories/2026/07/strike-pose-creating-more-realistic-multi-person-images