Ai scale image generation
How can I best make models like flash 3.6 image or gpt-image 2 get scale right from a 2d image. Use case is for swapping furniture in a room from a photo. It does extremely well in the realism and generation but sometimes it’ll alter the new table size to…
How can I best make models like flash 3.6 image or gpt-image 2 get scale right from a 2d image. Use case is for swapping furniture in a room from a photo. It does extremely well in the realism and generation but sometimes it’ll alter the new table size to make it fit where the old one went even if realistically it’s a much larger table. I can try giving it the table dimensions in the prompt but it’ll have no idea of room size. And if a user doesn’t want to have to Input the room plan on a website and just wants to upload a photo and have a “close enough” render in 30-45 seconds are there any work arounds like one size reference from a bed size or pot plant or something. Can the image generation models actually think that smart to take the table dimensions and the photo pot plant width and then scale it in the image?
Collected discussion
Moderator Announcement Read More » Hey u/Next-Customer-9613, If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt. If your post is a DALL-E 3 image post, please reply with the prompt used to make this image. Consider joining our public discord server! We have free bots with GPT-4 (with vision), image generators, and more! 🤖 Note: For any ChatGPT-related concerns, email support@openai.com - this subreddit is not part of OpenAI and is not a support channel. I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
It being such a frequent problem with image models, scaling always reliably, I run my prompt on two or three image models in use ai first to be sure of using the right one for my project.
It's not gonna reliably work. It can't even scale fully defined characters properly. There's a huge disconnect between the chat and image generation, so even if the conversation looks like it understands the spacial details the image generation isn't gonna work as well.