
January 22, 2025
Making AI Worlds
Spline is introducing Spell, an AI model to generate 3D worlds.
Spell is designed to generate entire 3D scenes or “Worlds” from an image, in just a few minutes. The worlds are consistent with the initial image input and are represented as a volume that can be rendered using Gaussian Splatting (or other methods, like NeRFs).

Image Input
Generated Video
Generated Volume
Capabilities
At its core, Spell is a type of diffusion model that can generate 3D worlds with realistic multi-view consistency across a wide range of categories, including people, objects, environments, 3D characters, and more.
The model can render images from multiple angles of a particular subject with high accuracy and detail, as well as generate controlled camera paths, all while staying consistent with the 3D scene.
Volumes generated by Spell
It is also capable of visually simulating physical material properties like reflections, refractions, surface roughness, and some camera properties like Depth of field, and even camera/object intersections when attempting to go inside surfaces.
Spell prioritizes physical consistency and aims to stay rooted in reality by simulating camera intersections with objects instead of interpolating/morphing to maintain visual flow. Example: If the camera gets inside a wall, it will simulate an actual intersection with the wall, instead of converting the wall into something else.
We found that the model can extrapolate knowledge very well across categories, in some cases, producing good results even in situations far from the initial data distribution.
