Strike a pose: Creating more realistic multi-person images
By Tom Fleischman, Cornell Chronicle
Generating an image of a person on a computer, using text prompts, is easy. Generating one with two people is similarly simple.
But creating an image of multiple people actually doing something, and faithfully reproducing not just the people but that thing they’re all doing? Not so easy.
A Cornell research group has devised a way to do just that, incorporating the poses of all individuals in an image to inform the generation, making the scene more believable and accurate.
“The current models for image generation can produce beautiful humans, but accurately capturing an interaction, and doing so in a diverse manner, is still a challenge,” said Wenxuan Peng, doctoral student in computer science, who presented “Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes” at the Association for Computing Machinery’s SIGGRAPH 2026 conference, July 19-23 in Los Angeles.
Co-authors are Hadar Averbuch-Elor, assistant professor of computer science at Cornell Tech and the Cornell Ann S. Bowers College of Computing and Information Science; and Bharath Hariharan, associate professor of computer science (Cornell Bowers).
The difference between their model and others, Averbuch-Elor said, is the intrinsic use of one person’s pose in generating the image. They also designed an iterative scene-construction scheme that progressively builds the scene one person at a time, using each component to inform the next.
“Because we’re doing it iteratively,” Averbuch-Elor said, “the first time we’re predicting only the first person’s pose, then this prediction is fed as a condition in order to predict the second person, and so on.”
Image-generation is prevalent in fashion and AI-based portrait synthesis, but current models typically focus on single-person generation and rely on the user providing accurate pose information, as opposed to incorporating a pose automatically.
Peng and the group used the image generator FLUX, developed and introduced by Black Forest Labs in 2024, as the backbone of their method. They trained their model on data from “Who’s Waldo,” a large-scale vision-language dataset developed by Averbuch-Elor and colleagues at Cornell and Cornell Tech in 2021 that contains thousands of images of interactions involving multiple individuals.
The researchers combined pose detection with a multimodal LLM to create aligned descriptions, poses and spatial regions for each person. The LLM is encouraged to arrange the per-person descriptions in a logical order, beginning with the primary actor, then those close to that person, followed by background or peripheral figures.
To iteratively compose an image, the first person is created; at each stage, the model predicts an updated pose map and image, and carries only the pose map forward to inform the next stage. One of their experimental prompts read: “In an AFL (Australian Rules Football) match, a player drives forward with the ball as an opponent battles to strip it, while a teammate races in for support.” From that, the model created a realistic three-person image.
To test their model for accuracy in terms of person-to-person interaction, Peng and the group developed the benchmark dataset “DrawWaldoWorlds,” which not only verifies the presence of multiple people, but also evaluates whether a model can correctly establish, as they write in the paper, “who does what to whom” by specifying both role-specific actions and person-to-person relations.
In multiple experiments plus a user study, the group’s method outperformed current programs in terms of faithful rendering of multi-person scenes. In the user study, involving 20 individuals, the group’s was preferred by an average of 2-to-1 over two versions of the FLUX image generator.
“It’s not that people disregarded pose in previous image-generation models,” Averbuch-Elor said. “But in existing methods, the user typically needs to provide the pose as input, and that takes time, especially if the interaction involves multiple people. In our work, the user does not need to provide that as input. The model internalizes this instead; this also allows for better capturing the interactions involving multiple people.”
Media Contact
Get Cornell news delivered right to your inbox.
Subscribe