🎙️AI 访谈库
WSJ 专访 Murati 谈 Sora:训练数据来自哪里?('我不确定是不是用了 YouTube 视频'名场面)
Mira Murati · 时任 OpenAI CTO

WSJ 专访 Murati 谈 Sora:训练数据来自哪里?('我不确定是不是用了 YouTube 视频'名场面)

OpenAI's Sora Made Me Crazy AI Videos—Then the CTO Answered (Most of) My Questions | WSJ

2024-03-13 · The Wall Street Journal (Joanna Stern) · 11m · 约 9 分钟读完 · 原文
Joanna Stern 与 Murati 逐帧审视 Sora 生成视频并追问发布时间与成本;当被问'训练数据是否包含 YouTube 视频'时 Murati 的支吾回答成为 AI 圈热议的经典时刻。

the video captures sort of the detail of the prompt when it comes to the hair and you know sort of like professionally styled uh women but you can also see some issues certainly especially when it comes to the hands these two women not real they were created by Sora open ai's text to video AI model but these two women very real I'm Mira morati C of open Ai and former CEO yes for two days in November when open aai CEO Sam Alman was momentarily ousted Mora stepped in now she's back to her previous job running all the tech at the company including Sora is our video generation model it is just based on a text prompt and it creates this hyper realistic beautiful highly detailed videos of one minute length I've been blown away by the AI generated videos yet also so concerned about their impact so I asked open aai to generate some new videos for me and sat down with Mora to get some answers how does Sora work it's fundamentally a diffusion model which is a type of generative model it creates a more distilled image starting from random noise okay here are the basics the AI model analyzed lots of videos and learn to identify objects and actions when given a text prompt it creates a scene by defining the timeline and adding detail to each frame what makes this AI video special compared to others is how smooth and realistic it looks if you think about film making people have to make sure that each frame continues into the next frame with the sense of consistency between objects and people and that's that's what gives you a sense of realism and sense of presence and if you break that between frames then you get this disconnected sense and reality is no longer there and so this is what Sora does really well you can see lots of that smoothness in the videos open AI generated from the prompts I provided but you can also see flaws and glitches a female video producer on a sidewalk in New York City holding a high-end Cinema Camera suddenly a robot Yanks the camera out of her hand so in this one you can see the M doesn't follow the prompt very closely the robot doesn't quite yank the camera out of her hand but the person sort of morphs into the robot um yeah a lot of imperfections still one thing I noticed there too is when the cars are going by they change colors mhm yeah so while the model is quite good at continuity it's not perfect so you kind of see the Yellow Cab disappearing from the frame there for a while and then it comes back in a different frame would there be a way after the fact to say fix the taxi cabs in the back yeah So eventually that's what we're trying to figure out um how to use this technology as a tool that people can edit and create with I wanted to go through one other what do you think the prompt was uh it looks like the bull in a china shop yeah metaphorically you know you'd imagine everything breaking in the scene right and you see in some cases that uh the bull is stomping on things and they're they're still perfect they're not breaking that's to be expected this early on and eventually there's going to be more stability and control and more accuracy in reflecting the intent of what you w and then there was that video of well us the woman on the left looks like she has maybe like 15 fingers in one of the shots hands actually have their own way of of motion and it's very difficult to simulate the motion of of hands in the clip the mouths move but there's no sound so is audio something you're working on with Sora with Sora speciic specifically not in this moment but we will eventually every time I watch a sort clip I wonder what videos did this AI model learn from did the model see any clips of Ferdinand to know what a bull in a china shop should look like was it a fan of SpongeBob wow you look real good with a mustache Mr Krabs by the way my prompt for this crab said nothing about a mustache what data was used to train Sora we used publicly available data and licensed data so videos on YouTube I'm actually not sure about that okay videos from Facebook Instagram you know if they were publicly available um available yeah publicly available to use um there might be that data but um I'm I'm not sure I'm not confident about it what about shutter stock I know you guys have a deal with that I'm I'm just not going to go into the details of of the data that was that was used but it was publicly available or licensed data after the interview morani confirmed that the licensed data does include content from Shutterstock those videos are 720p 20 seconds long how long does it take to generate those it could take a few minutes depending on the complexity of the prompt our goal was to really focus on developing the best capability and now we will start looking into optimizing the technology so people can use it at low cost and make it easy to use to create these you must be using a lot of computing power can you give me a sense of how much computing power to create something like that versus a chat gbt response or a do image chat gbt and do are optimized for the public to be using them whereas Sor is really a research output it's much much more expensive we don't know what it's going to look like exactly when we make it available eventually to the public but we're trying to make it available at similar cost eventually to what we saw with Del you said eventually when is eventually I'm hoping yeah definitely this year but could be a few months there's an election in November you think before or after that you know that's suddenly a consideration dealing with the issues of um misinformation and Har ful bias and we will not be releasing anything that we don't feel confident on when it comes to how it might affect global elections um or yeah other other issues right now Sora is going through red teaming aka the process where people test the tool to make sure it's safe secure and reliable the goal is to identify vulnerabilities biases and other harmful issues what are things that just you won't be able to generate with this well we haven't made those decisions yet but I think there will be consistency on our platform so similarly to Del where you can't generate uh images of public figures I expect that we will have a similar policy for Sora and right now we're in discovery mode and we haven't figured out exactly where all the limitations are and how we'll navigate our way around them what about nudity I'm not sure you can you can imagine that you know there are creative settings in which artists might want to have more control over that and right now we are working with artists and creators from different fields to figure out exactly what's useful what level of flexibility should should the tool provide how do you make sure the people who are testing these products aren't being inundated with elicit or harmful content that's that's that's certainly difficult and in the very early stages it is part of red teaming um something that you have to take into account and make sure that people are willing and able to do it um when we work with contractors we go much further into that process but that is certainly something difficult we're laughing at some of these videos right now but people in the video industry may not be laughing in a few years when this type of technology is impacting their jobs you know the way that I see it is this is a tool for extending creativity and we want people in the film industry creators everywhere to be a part of informing how we develop it further and also how we deploy it and also you know what are the economics around using these models when people are contributing data and such one thing was clear from all this this Tech is going to quickly get faster better and become widely available how are we going to tell the difference between what is real video and what is AI video we're doing research in water marking the videos but really figuring out content Provence and how do you trust what is real content versus something that happened in reality versus you know content created for misinformation and this is the reason why we're actually not deploying the systems yet because we need to figure out these issues before we can confid L deploy them broadly that was reassuring to hear but there are still big concerns about silicon valleys raised to create AI tools and its ambition for power and money versus our safety it's not really a difficult demand or a difficult balance between profit and safety guard rails I'd say the hard part is really figuring out the safety questions and the societal questions that that's that's really what keeps me up at night there's this amazement about the product but then we've also talked about all of these concerns is it worth it it's definitely worth it AI tools will extend our creativity and knowledge Collective imagination ability to do anything it's going to be extremely hard along the way to figure out the right path to Bringing AI tools into our day-to-day reality but I think it's definitely worth trying