Sora, February 2024: OpenAI Text-to-Video Preview

On this page6 sections

Sora, February 2024: OpenAI Text-to-Video Preview

Sora Preview
Date:
February 15, 2024 (preview) · December 9, 2024 (public release)
Lab/Organisation:
OpenAI
Paper/Outcome:
“Sora: Creating video from text” (blog) + “Video generation models as world simulators” (technical report)
Significance:
Preview that signaled AI video was the next multimodal frontier
Sora

OpenAI’s text-to-video model, previewed in February 2024. Sora produced minute-long, physically-plausible videos from text prompts, and was the first generative model whose outputs were widely mistaken for actual video footage. Sora remains in limited release.


The Announcement

Sora is OpenAI’s text-to-video model, previewed on February 15, 2024. It produces minute-long, physically-plausible videos from text prompts, and was the first generative model whose outputs were widely mistaken for actual video footage. Sora remains in limited release as of 2025.

The February 15, 2024 announcement consisted of two documents. The first was a public blog post, titled “Sora: Creating video from text,” which introduced Sora to a general audience and which included a gallery of sample videos. The second was a technical report, titled “Video generation models as world simulators,” which provided a high-level description of the model’s architecture and capabilities. The technical report was, by the standards of AI research, unusual — it described the model’s approach but it did not include the full technical details, the training data, or the code. It was, in effect, a disclosure without a release — a way of telling the world what OpenAI had built without giving the world the ability to build it themselves.

The sample videos, which accompanied the blog post, were the announcement’s most striking feature. The videos — generated from text prompts — were high-definition, up to sixty seconds long, and they depicted scenes of remarkable complexity and visual fidelity. A prompt like “a stylish woman walks down a Tokyo street filled with warm glowing neon and animated city signage” produced a video that showed exactly that — a woman walking, the neon signs reflecting in puddles, the camera tracking her movement — with a level of detail and coherence that no previous AI video generator had achieved. The videos were not perfect — they contained artefacts, physics errors, and moments of visual incoherence — but they were, by most accounts, dramatically better than anything the public had seen before.

The announcement was, by most measures, a viral sensation. The sample videos were shared millions of times on social media. They were covered by every major news outlet. They became, in the days following the announcement, the dominant topic of conversation in the AI community, in the technology industry, and in the broader public. The “Sora moment” — the phrase that came to be used for the cultural impact of the announcement — was, by most accounts, the moment when AI-generated video crossed a threshold, from a niche research curiosity to a mass-market phenomenon.


The Technology

The technical details that OpenAI disclosed, while not complete, were sufficient to give a sense of the model’s approach. Sora, the technical report explained, was a diffusion transformer — a diffusion model (the same family as Stable Diffusion, see B78) built on a transformer architecture (the same family as GPT, see B70). The model operated on “spacetime patches” — small, spatiotemporal units of video and image data — that allowed it to handle the temporal dimension of video in a unified way. The approach was, in some ways, a natural extension of the diffusion-model architecture that had powered the image-generation revolution to the video domain. But the extension was, by most accounts, non-trivial — video, with its temporal dimension, its coherence requirements, and its much larger data volume, posed challenges that image generation did not.

Why “spacetime patches” mattered

The “spacetime patch” was Sora’s key architectural innovation. Image diffusion models operate on spatial patches — small grids of pixels. Video adds a temporal dimension: a video is a sequence of frames, each a spatial grid, and the model has to maintain coherence across the sequence. Sora’s solution was to treat space and time uniformly: instead of separate spatial patches per frame, the model used “spacetime patches” — small cubes of spacetime that captured both the spatial content of a frame and the temporal relationship to adjacent frames. This let the model handle video with the same machinery it used for images, and it let it generate videos of varying resolutions, aspect ratios, and durations without retraining. The “world simulator” framing in the technical report’s title reflected OpenAI’s claim that this unified spacetime representation was, in effect, learning a model of how the physical world behaves — not just pixel patterns, but object permanence, gravity, and physical causality. (Whether the model was actually “simulating the world” or merely producing plausible-looking video was, however, a matter of substantial debate.)

The technical report described the model’s capabilities — videos up to sixty seconds, high-definition, with complex camera motion and detailed scenes. But it did not disclose the training data, the model size, the compute used, or the specific architectural details. The omission was, by most accounts, deliberate. OpenAI, which had been moving toward a more closed approach in the preceding years, was not interested in giving competitors — or potential misusers — the information they would need to replicate the model. The decision was, within the AI research community, controversial. Many researchers argued that the lack of disclosure violated the norms of scientific openness, and that it made it impossible to independently evaluate the model’s capabilities, its limitations, or its risks. OpenAI defended the decision on safety grounds, arguing that the risks of full disclosure — of giving bad actors the information they would need to build their own video-generation systems — outweighed the benefits.


The Non-Release

The most striking aspect of the Sora announcement was the non-release. OpenAI did not make Sora available to the public. The model was given only to a small group of “red teamers” — safety testers who were tasked with probing the model for vulnerabilities, biases, and potential misuses — and to a selected group of visual artists, designers, and filmmakers, who were asked to provide feedback on the model’s creative capabilities. The general public, despite the viral attention that the announcement generated, could not use Sora.

The non-release was, by OpenAI’s own account, a deliberate choice. The company cited several concerns. The first was the risk of deepfakes — AI-generated videos that could be used to create false or misleading content, particularly of public figures. The second was the risk of misinformation — AI-generated videos that could be used to spread false information, particularly in the context of an election. The third was election integrity — 2024 was a US presidential election year, and the prospect of AI-generated videos being used to influence the election was, for OpenAI and for regulators, a serious concern.

The concerns were, by most accounts, legitimate. AI-generated video, if widely available, could be used to create deepfakes at a scale and quality that would be difficult to detect and counter. The prospect of such deepfakes being used in an election — to create false videos of candidates, to spread fabricated news, to manipulate public opinion — was, for many observers, one of the most alarming scenarios of the generative AI era (see B84). OpenAI’s decision to withhold Sora, at least until the election was over, was, in this context, a cautious and responsible choice.

The non-release was, however, also controversial. Critics argued that OpenAI was using safety concerns as a cover for a broader strategy of closed development — that the company was withholding the model not primarily for safety reasons but to maintain its competitive advantage. The critics noted that OpenAI had, in its early years, been a champion of openness — the company’s original mission had been to build AGI for the benefit of humanity, and the “Open” in its name had been a reference to open source. The company had, over the years, moved steadily away from that commitment, becoming one of the most closed and secretive of the major AI labs. The Sora non-release, the critics argued, was another step in that direction.


The Competitive Landscape

Sora was not, of course, the first AI video generation model. The field had been developing for several years, and there were several competitors. Runway, the AI company that had co-developed the latent diffusion architecture (see B78), had released its Gen-2 video generation model — research announced in February 2023, introduced in March 2023, and made publicly available in June 2023. Pika, another AI video start-up, had launched Pika 1.0 on November 28, 2023 (alongside $55 million in funding). Google had announced Lumiere, its video generation model, in January 2024, just weeks before the Sora announcement. Meta had developed Emu, its own video generation model. And Stable Video Diffusion, an open-weights model from Stability AI, had been released as a research preview on November 21, 2023.

The competitive landscape was, in some ways, a parallel to the image generation landscape that had existed before Stable Diffusion. There were closed, API-based services (Runway, Pika), open-weights models (Stable Video Diffusion), and research announcements from the large labs (Google, Meta). Sora, when it was announced, was the most capable of these — the sample videos were, by most accounts, the most impressive — but it was not, in its approach, fundamentally different from its competitors. The difference was in the quality, the resources, and the attention that OpenAI could bring to bear.

The competitive landscape also, in some ways, explained OpenAI’s non-release strategy. If Sora had been released, it would have been, given OpenAI’s resources and brand, the dominant video generation model. The competitors — Runway, Pika, the open-weights models — would have been marginalised. By withholding Sora, OpenAI gave the competitors a window — a chance to develop their own models, to build their own user bases, and to establish their own positions in the market. The non-release was, in this sense, a competitive concession as well as a safety measure.


The Actual Release

The actual public release of Sora came on December 9, 2024 — nearly ten months after the February preview, and as part of OpenAI’s “12 Days of OpenAI” livestream event (which ran daily from December 5 through December 20, 2024; Sora was unveiled on Day 3). The release made Sora available as a standalone product at Sora.com, branded as “Sora Turbo,” to ChatGPT Plus ($20/month) and Pro ($200/month) subscribers in the United States and Canada. The release was, by most accounts, more limited than the preview had suggested — the publicly available version of Sora had some constraints on video length, resolution, and usage that the preview videos had not shown. But the release did make Sora, for the first time, available to the general public.

The ten-month gap between the preview and the release was, by most accounts, longer than the AI community had expected. The gap was, in part, a consequence of the safety concerns that had motivated the non-release — OpenAI spent the months between the preview and the release conducting safety testing, engaging with regulators, and developing safeguards. But the gap was also, in part, a consequence of the technical challenges of deploying a video generation model at scale — the compute requirements, the latency, the user interface, and the integration with ChatGPT all required substantial engineering work.

The release, when it came, was met with a mix of excitement and disappointment. The excitement was about the capabilities — Sora, even in its constrained public version, was the most capable video generation model available to the public. The disappointment was about the constraints — the limits on video length, the limits on resolution, the limits on usage — and about the fact that the public version was, in some ways, less capable than the preview videos had suggested. The gap between the preview and the release was, for some users, a reminder of the difference between a curated demonstration and a deployed product.


The Legacy

The legacy of the Sora preview is, by any measure, the legacy of AI-generated video as a cultural and technological phenomenon. The February 15, 2024 announcement was the moment when AI-generated video crossed a quality threshold — when the videos became good enough to be mistaken, at least briefly, for real footage, and when the implications of that capability became impossible to ignore. The “Sora moment” — the cultural impact of the preview — shaped the public conversation about AI video, about deepfakes, about election integrity, and about the future of media.

The legacy is also, in important ways, the legacy of the non-release strategy. OpenAI’s decision to preview Sora without releasing it was, by most accounts, a new approach in the AI industry. Previous major AI announcements — GPT-3, DALL-E 2, ChatGPT — had been accompanied by releases, either through APIs or through public products. Sora was the first major AI announcement that was, by design, a preview without a release. The approach was, for the industry, a signal — a signal that the major AI labs were becoming more cautious about the deployment of their most capable models, and that the era of “move fast and break things” was, at least for the frontier labs, giving way to a more deliberate, more safety-conscious approach.

The legacy is also, finally, a reminder of the trade-offs of the AI era. Sora’s capabilities — the ability to generate high-quality video from text — were, by most accounts, genuinely impressive, and they promised a range of beneficial applications, from filmmaking to education to communication. But the same capabilities — the ability to generate realistic video — also posed serious risks, from deepfakes to misinformation to the erosion of trust in visual media. The trade-off — between the benefits and the risks, between openness and safety, between innovation and caution — was, in the Sora preview, made vivid. The field has not, as of 2026, resolved the trade-off. But the Sora preview was the moment when the trade-off became, for the broader public, impossible to ignore.


Further reading

Series Companions

This piece is part of Minds & Machines: Beyond the Series. The companion pieces B78 — Stable Diffusion Release, August 2022 (the image-generation model whose diffusion architecture Sora extended to video), B77 — The Colorado State Fair AI-Art Win, August 2022 (the cultural moment that previewed the AI-art controversies Sora would intensify), B27 — GAN, June 2014 (the earlier generative architecture that diffusion models replaced), and B84 — The AI Election Revisited (the election-integrity context that motivated the Sora non-release) cover the related milestones.

The legacy of Sora is still being written. The decisions made today — by engineers, by policymakers, by users — will determine whether it is remembered as a turning point or a footnote.