← Back to LAB

THE AGENTIC REVOLUTION

Runs September 7, 2026 14 min read Last updated September 11, 2026
The Cinema 4D interface with the chase scene open, the object manager on the right listing the camera pivot, the two car rigs and the track.
The Cinema 4D viewport, with the scene the MCP built.

I wanted to see if I could drive Cinema 4D through an MCP server, with no hand on the mouse. The basic shapes it produced were then handed to an image model and a video model. The interesting part was watching a really ugly previs (but super fast to make) turn into a beautiful render in minutes, while I kept control.

The idea

I had seen plenty of people driving Blender from a chat window through an MCP server, and I wondered if Cinema 4D could do the same. 2 questions, chained.

  • Can a 3D application be driven end to end by a model through a tool protocol: build the scene, rig it, light it, animate it, render it?
  • Is the rough output of that scene worth anything as an input to generative video?

I tested 4 different scenes to see what it could do, to push the ideas and the techniques, and to see what each step cost as of today.

MCP server

It is a strange feeling to watch a scene being built from a text prompt. Something magical and scary at the same time. Under the hood it used Python to create everything: the objects, the connections, the relationships.

Across the scenes the tool calls ran between 20 and 63 per scene, and not one construction call failed. Objects, hierarchy, parameters, cameras, materials, rigs, render settings and saving all held. The detail of what worked and what bit is at the end of the page, after the shots it produced.

Project 01The race track

I built 4 different environments from one blocking: dry asphalt in daylight, snow, a cyberpunk neon night, and a sand track by the sea.

Step 1Block the shot

What the tool did not settle was the car animation (too basic). I adjusted the final camera framing and the placement of the 2 cars by hand. The MCP carried the scene to the point where a taste decision took a minute, and no further.

The blocked, coloured cars that drove the idea: the red one starts behind, then goes in front.

Step 2Frame zero to plate

A single frame of the blocking went to many image models, which returned the look of the finished shot. The camera, the crop and the horizon stayed where the previs put them, because everything downstream follows them.

Swap between models and looks in the player below, and drag across the picture to check each one against the grey plate it came from.

0:00 / 0:00
Opacity 70%

What the frames cost?

The cheapest lever in the whole chain (by far). On this shot I generated 15 frames across 4 models, all sent to Magnific, $3.88 in total, most of it on the 2K GPT 2 frames for the weather changes.

THE FRAMES COST#01-chase15 frames$3.88
$0.42
Seedream 5 Pro2560 × 14405 frames
$0.13
Nano Banana Pro2752 × 15362 frames
$0.13
Nano Banana 22752 × 15362 frames
$3.20
GPT 21344 × 752 · 4 in 2K6 frames

The power here is to iterate!

4 looks came out of 4 starting frames with nothing else changed, for less than the price of a coffee. Every decision that can be made on a still should be made before a single second of video is generated. I did not push the resolution to 4K, which I would recommend to get the most out of it.

Step 3Write the prompt

That part is made for an LLM. I could write it myself, but it is tedious (and expensive down the line when it goes wrong). It is harder than it sounds. What a model is actually good for here is not the prose. It is memory of the reasons.

Prompt sent to Seedance 2.0 Drag the corner, or open it in full
The reference image is the exact start frame and the visual style anchor. Frame 0 must match it exactly: same two cars, same circuit, same camera angle, same lens, same crop, same horizon line, same daylight, same composition, same overall look.

The reference video is the motion source for the camera path, framing and timing. Match its camera movement as closely as possible across the full eight seconds. Do not reinterpret the camera path, do not shorten it.

The action, and this is the critical point: this is a lateral tracking shot. The camera travels alongside the two cars at a constant distance for the whole shot, neither closing in nor falling back. The cars sweep continuously through a long left-hand bend with their noses angled toward the lower right of frame, seen from a high rear three-quarter angle. They are neither driving away from the lens nor toward it. The circuit, the kerb line and the tyre barriers stream past behind and beside them, which is what carries the speed. The red car is ahead, the blue car behind it, and the gap between them closes over the shot.

Keep both cars locked throughout. No redesign, no shape drift, no identity drift, no material drift, no logos, no badges, no text, no branding. The nearer car stays blue, the further car stays red. Both cars stay on the asphalt inside the white kerb line and never cut across the run-off area.

Preserve the environment of the reference image throughout: same ground, same horizon, same daylight, same grade.

Smooth, photoreal, grounded, temporally consistent across the full clip.

Step 4The fun part! Image + previs

0:00 / 0:00
Opacity 70%

4 worlds, same timing. The previs carried that. The last one is a power move, right? :)
Anyhow, the magic happened!

What the videos cost?

Video is where the money went, and the gap is not subtle. On this shot I ran 13 video takes for $43.20: 9 rehearsals on Seedance 2.0 at 720p and 1 take on Seedance 2.5 at 1080p, sent to Magnific, and 2 takes sent to fal at 1080p, both with the cars driving the wrong way.

THE VIDEOS COST#01-chase13 takes$43.20
$17.14
Seedance 2.0720p9 takes
$7.34
Seedance 2.51080p1 take
$18.72
Seedance 2.5 on fal1080p2 takes

Project 02City drone shot

Step 1Block the city

The MCP built the city in a few minutes: a grid of blocks with a few variations, a ground plane with a hard edge, a very basic setup. I animated the camera, starting high and dropping toward a junction. This was the first version, a straight descent with some banking. This render showed the UI, which eventually made it into the generated video (oops, my bad).

The first blocking, looping: a straight drop into the grid.

A straight drop was an easy shot, and easy shots do not test a video model. The point of this project was a harder one: a descent that turns as it falls, a full circle around the block before landing. It is a difficult move (most video generations lost the geometry at the fastest point of the rotation). This time I cleaned up the UI and made sure only the needed shapes were kept.

The second blocking: a full 360, with banking.

Step 2Find the right image

This one brought a lot of satisfaction! I mean, the previs was just cubes and rectangles on a grid! Very blunt. A few minutes later the blocks became glass, haze and rooftop skyscrapers, with so much detail. A little input, a lot of picture. Finding the right one was still a search, not a click. The same plate went to the same image models with the same prompt, word for word, including an instruction to erase every trace of the 3D interface.

0:00 / 0:00
Opacity 70%

What the frames cost?

8 frames on this shot, $0.67, for a city that took minutes to block.

THE FRAMES COST#02-drone8 frames$0.67
$0.26
Seedream 5 Pro2560 × 14403 frames
$0.13
Nano Banana Pro2752 × 15362 frames
$0.06
Nano Banana 22752 × 15361 frame
$0.22
GPT 21344 × 7522 frames

Step 3The prompt

The prompt had the same shape as the one in project 01: the reference image is the start frame, the reference video is the motion, and the camera is described as what it is already doing. What this one added was traffic, asking for a city alive in the 360. Here is the exchange around it, and what it took to get the density I asked for.

C Claude #02-drone-descent
Mr. Pinoux
Two drones in 1080p, with traffic. Cars on the avenues, people on the crossings. I want the junction busy when the camera lands.
Claude

Both are queued, 5,600 credits each. The instruction went in at length: cars, taxis, vans and buses on every avenue in both directions, queues building and releasing at the intersections, tiny pedestrians on the marked crossings. Traffic follows the grid exactly as it is drawn in the plate, never over a roof.

Honest verdict: the instruction went through, but not at the density asked. The descent is clean and lands at street level with vehicles moving, crossings, silhouettes. It is alive, but it is a Sunday morning, not rush hour. On the GPT 2 frame the massing also drifts in the last third: the block the camera lands on is no longer the one in the plate. The Seedream frame holds better. If you want density, say it bluntly, bumper to bumper, stopped queues, rush hour. Same price to retry.

Mr. Pinoux
Before rewriting anything, look at where the 2.0 lands. The previs drops into a tight canyon and the video lands on a wide avenue that is not there. The geometry lets go when the rotation gets complex. Try the same thing on the 2.5.
Claude

Your read is right, and there is a bonus. On geometry: the previs lands in a tight canyon, two parallel walls, the street running away in the centre. The 2.0 lands on a wide avenue that is not the previs's, and the buildings on the left changed height. The 2.5 finds the canyon at the right proportions, and during the rotation itself, at frames 120 and 150, the 2.0 is already smearing while the 2.5 holds a coherent city.

The bonus: the traffic density arrives. Same previs, same still, same prompt word for word, only the model changes. Full avenue, yellow taxis, vans, pedestrians on the pavements. That is exactly what was asked of the 2.0 and what it would not give. 8,640 credits.

@Claude and now the same eight seconds on the newer model

The text below is the traffic prompt, exactly as I sent it.

Prompt sent to Seedance 2.0, then to Seedance 2.5 unchanged Drag the corner, or open it in full
Use the still image as the locked look of this shot: same city, same building massing, same street grid, same facades, same rooftops, same light direction, same haze.

Use the reference video only as the motion source for the CAMERA. Follow its camera path, its timing and its speed exactly, as one continuous take: a drone descending over the city. Do not invent a different camera move. Do not cut.

THE CITY IS ALIVE. Traffic runs continuously through the whole shot:

- Cars, taxis, vans and buses flow along every avenue and cross street, in both directions, keeping to their lanes. Vehicles are small at this altitude but clearly moving, not parked.
- Traffic follows the street grid exactly as it is drawn in the plate: it never drives over a building, never crosses a roof, never invents a new road.
- Queues build and release at intersections, a few vehicles turn at corners, gaps open and close between them.
- Tiny pedestrians move along the pavements and across the marked crossings.
- Windows catch and lose reflections as the camera descends, rooftop vents and fans turn slowly, a flag moves, a construction crane rotates gently if one is visible.
- Light haze drifts between the towers, and shadows stay exactly where the still puts them.

Every building keeps its exact footprint, height and position. Do not add or remove buildings. Do not change the street layout. No camera cuts, no speed ramp, no time-lapse: this is real time.

Photoreal aerial cinematography, natural daylight, fine grain. No text, no logos, no branding.

Step 4The fun part! Image + previs

2 previs, so 2 videos. The first one is the simple landing. The second one is the 360.
A clean descent, and traffic moving at street level when the camera arrived. This first one worked pretty well from the get-go, but it felt like a very quiet city, too quiet.

0:00 / 0:00
Opacity 70%

The 360. Updated previs, new still to match, updated prompt with more traffic and more life in the city. Seedance 2.5 found the canyon at the right proportions, some pops, with a giant roof on the floor (overall it still felt better). Seedance 2.0 could NOT hold the scene at the high stress point, when the camera is upside down. I kept those in the rejected videos down below. Seedance 2.5 handled it more elegantly. Still, I feel the fast camera change was a bit hard for the model.

0:00 / 0:00
Opacity 70%

This one was pricey! The way I spent that money was in 2 stages: check the prompt on the 720p Seedance 2.0, where a take costs $1.90 and shows in a minute whether the camera holds and the words are understood, then run the 2.5 at 1080p once there was nothing left to change. On this shot: 9 takes, $41.33.

THE VIDEOS COST#02-drone9 takes$41.33
$7.60
Seedance 2.0720p4 takes
$9.52
Seedance 2.01080p2 takes
$9.52
Seedance 2.04K1 take
$14.69
Seedance 2.51080p2 takes

Project 03The still life

The product pack shot. A product, a small camera move (pretty basic).
Here I went with changing the objects into something else entirely, which is easy for the image generation models.

Step 1Block the shot

What came out of the MCP was blocky and dull: a bottle, a wide cylinder, 2 spheres, a ring and a disc on a chamfered slab, lit 3-point (again, the idea here was speed).
I added a simple orbital camera move.

4 seconds of previs, looping: 6 primitives on a rectangle and a simple orbiting camera.

Step 21 plate, 3 worlds

This is where the ideas could be as wild as I wanted: A realistic still life? Sure, check. What about a miniature futuristic city? Why not. What about a metallic forged citadel riding a lake of lava? Come on, let's gooo!

0:00 / 0:00
Opacity 70%

Lock the shot, iterate around it for next to nothing, and be as creative as the brief allows.

THE FRAMES COST#03-still-life16 frames$1.28
$0.34
Seedream 5 Pro2560 × 14404 frames
$0.26
Nano Banana Pro2752 × 15364 frames
$0.26
Nano Banana 22752 × 15364 frames
$0.42
GPT 21344 × 7524 frames

Step 3Writing the prompt

The lava is where I rewrote the prompt, twice. The first instruction produced a slab floating above the lava, with the lava reading as a cracked crust rather than as a liquid.

C Claude #03-still-life
Mr. Pinoux
ok, so for this one - can we get the lava river to flow right to left, interacting with the metal plate, showing the cooling down as the lava touches the metal plates, but always in motion. No slow motion please
Claude

Two versions, on the two frames you kept. The prompt now says three things the old one did not. The direction, in plain words: right to left in screen space, whatever the camera does. The speed: real time, no slow motion, no ramp. And the cooling cycle at the contact: white at the point of impact, then yellow, orange and a black crust as the flow moves away.

On the stills the thermal cycle already reads: the rim goes from bright orange to black crust in four seconds. The direction only shows in motion, so that one is yours to judge in the comparator.

@Claude same instruction on the Nano Banana 2 frame

Here is the prompt that came out of it, exactly as I sent it.

The third lava prompt, sent to Seedance 2.0 Drag the corner, or open it in full
Use the still image as the locked look of this shot: same forged iron raft, same blackened tower, same cast iron drum, same three hammered orbs, same iron plates, same riveted cupola, same torus, same night, same grade.

Use the reference video only as the motion source for the CAMERA. Follow its camera path, its timing and its speed exactly, as one continuous take. Do not invent a different camera move. Do not cut.

SPEED. This is real time at normal speed. No slow motion, no speed ramp, no time-lapse, no dreamy floating. The lava behaves like a fast, hot, runny river, not like thick syrup.

DIRECTION. The lava river runs from the RIGHT edge of the frame toward the LEFT edge, and it keeps that direction for the entire shot whatever the camera does. Read it in screen space: anything floating on the surface enters from the right and exits on the left. In four seconds a given patch of crust must travel at least half the width of the frame. If the flow drifts the other way, or stalls, the result is wrong.

CONTACT AND COOLING. The river hits the raft on its RIGHT side and splits around it, closing again in a wake on the LEFT side. At the point of impact the metal is white hot. As the lava sweeps past and away along the flanks, that same metal cools in front of the camera: white to yellow to deep orange to a dull grey-black skin, and a thin crust forms on the iron and then cracks and flakes off as the next wave hits. That cooling cycle repeats several times during the shot, faster on the right where the impact is, slower along the left flank in the wake.

ALWAYS MOVING. The dark surface crust is broken into plates that ride the current from right to left, with brilliant molten cracks opening and closing between them. Waves slap the raft edge and throw spatter and sparks. Glowing runnels drip off the chamfered edge and are immediately carried away to the left. Heat haze shears upward and drifts. Embers rise.

The forged objects do not move, do not melt and do not change shape. The raft holds its position, riding the flow with only the faintest settle. Every object keeps its exact place and silhouette.

Photoreal, cinematic night, volumetric heat, fine metal texture, real-time speed. No text, no logos, no people, no hands.

Step 4The fun part! Image + previs

The still life, the miniature city, the lava. The camera move is identical in every generation.

0:00 / 0:00
Opacity 70%

What the videos cost?

15 takes on this shot, $18.43, most of them 4-second rehearsals on the 2.0 at 720p.

THE VIDEOS COST#03-still-life15 takes$18.43
$12.38
Seedance 2.0720p13 takes
$2.38
Seedance 2.01080p1 take
$3.67
Seedance 2.51080p1 take

Project 04The beach

On this shot I had to try to bring a character in.

Step 1Block the hut

A building, a basic palm, 2 shapes for the beach and the water, 1 crude character with a super basic animation. That was the whole previs for 7 seconds. Again, everything but the camera animation was made by the MCP. I added the colours and fixed the arm animation, which was flipped.

7 seconds of previs. The hut, the palm, the figure stepping out, the camera pushing in.

Step 2Character sheet + location

A character sheet, 3 views of one woman.

Three views of the same woman, the character sheet passed as a reference to every take.
generated with Seedream 5 Pro and passed as a character reference.

The location. It took several plates to get a first frame that worked, because what the previs shows is what the model is given (I think it was missing context).

The first plate was plain grey shading, and the frame that came back was handsome but not aligned on the camera: no sea, no horizon, a beach invented from nothing.

The second switched to Gouraud shading with lines and a grid on the ground, which gives a previs more structure to read. It gave the model the same lines, and GPT 2 rendered them as lines in the sand, tracks and ruts running across the beach.

Back to Cinema 4D, then. I placed simple coloured shaders on the ground and the water: a bit of shade, no hard lines on the sand or the sea. The lines stayed on the hut and on the figure.

On the figure it made no difference, since the figure was replaced entirely. On the hut it did: look at the right side of the roof in the last frame, where the shading edge of the previs became an artefact in the corrugated iron. Not worth more time on a test, and worth knowing.

0:00 / 0:00
Opacity 70%

What the frames cost?

17 frames on this shot, $1.47, the most of the 4 projects, because the location took 4 plates to get right.

THE FRAMES COST#04-beach17 frames$1.47
$0.51
Seedream 5 Pro2560 × 14405 frames
$0.26
Nano Banana Pro2752 × 15364 frames
$0.26
Nano Banana 22752 × 15364 frames
$0.44
GPT 21344 × 7524 frames

Step 3The prompt takes the camera

Pass 1: no interior described. It tracked the previs: the whole hut in shot, sea and horizon on both sides, raised arms complete.

Pass 2: I added a paragraph describing what was inside the hut, and the previs wasn't the only source of truth anymore, and I lost control with it.

Identical at the first frame, tighter by the middle. The location was traded away. Here is the exchange that found out why.

C Claude #04-beach
Mr. Pinoux
The framing drifts on the beach shot. Same previs, same starting frame, same model as the pass that worked. Only the text changed. Why.
Claude

My mistake, and it is in the text I wrote. The paragraph describing the interior contains as the camera comes closer. That is a camera instruction, and it competes with the reference video for the same thing. The reference lost.

There is a second cause that is not a writing slip. The prop list asks for a paperback left face down and small jars. Those only read at a certain size, so asking for a detail to be legible is asking for a shot size, even with no camera verb anywhere.

Mr. Pinoux
So rewrite it. I want the interior, and I want the previs framing back.
Claude

Three changes. Every camera verb removed. The interior described as a state rather than as a reveal, so nothing implies a move toward it. The prop list cut down to what reads from the reference distance.

Then an explicit lock as its own paragraph: the camera never comes closer than the reference, and whatever is too small to read stays too small. That last clause is the one that holds, because it gives the model a way to obey the prop list without moving.

@Claude rewrite it with the interior and the framing locked

Pass 3 is the one below, entire. The lock is the third paragraph.

Prompt of pass three, sent to Seedance 2.5 Drag the corner, or open it in full
Use the still image as the locked look of this shot: same beach hut, same coconut palm, same sand, same sea and horizon, same warm grade.

Use the reference video as the motion source AND as the framing reference. Follow its camera path, its timing, its speed and its shot size exactly, as a single continuous take. Do not invent a different camera move. Do not cut.

FRAMING LOCK. The camera never comes closer to the hut than the reference video brings it. At every moment the hut occupies the same portion of the frame as in the reference. At the last frame the whole hut is still inside the frame, with the sea and the horizon visible on both sides of it, exactly as in the reference. If the hut fills the frame, if the walls run off the edges, or if the sea and the horizon leave the picture, the result is wrong.

The woman in the character reference is the person in this shot. Keep her exactly as she appears there: same face, same hair, same build, same dress and sandals. She steps out of the doorway at the moment given by the reference video, walks forward a few paces along the sand, and raises both arms overhead in a wide stretch at the very end, exactly on the reference timing. Her raised hands stay inside the frame.

The hut is lived in, and the open doorway shows it at whatever size that doorway happens to be. Inside: a made bed with white linen against the far wall, a shelf with a kettle and enamel mugs, a jug of flowers, a woven rug on the floorboards. A warm low lamp burns in there, so the doorway glows against the cool blue of the sea outside and reads as a warm pocket rather than a black hole. Do not move the camera to make any of this legible. Whatever is too small to read at the reference distance simply stays too small to read.

Palm fronds and dune grass move in a light onshore breeze. Long warm shadows travel across the sand as the shot advances. Small waves break on the shoreline behind.

No text, no logos, no on-screen graphics, no other figures.

Step 43 takes, and 1 to avoid

First, the pass with nobody in it, Seedance 2.0, which showed the problem: the character was refused (3 times out of 3!) and the shot came back empty.

Second time, the pass with the character on Seedance 2.5, which accepted her.

Third time, I described the interior in too much detail, and the camera drifted away from the previs toward the end.

The fourth time, with less detail on the interior, the framing locked. That was the pass that closed the project.

0:00 / 0:00
Opacity 70%

This shot gave me a new set of surprises. As it was the only one with a character, I learnt the hard way: refused 3 times out of 3 by Seedance 2.0, with a clothing description, without it, and finally with no description at all but the character sheet as reference. Tried fal, Magnific, nope. Maybe that would have changed with a Chinese platform, but who cares.

Only Seedance 2.5 was able to use her without problems.

0:00 / 0:00
Opacity 70%

What the videos cost?

8 takes on this shot, $27.13. The 2.5 cost the most, and it was the only one that accepted the character.

THE VIDEOS COST#04-beach8 takes$27.13
$5.00
Seedance 2.0720p3 takes
$9.28
Seedance 2.5720p3 takes
$12.85
Seedance 2.51080p2 takes

What held

4 projects, 1 day, and 4 things that held every time I tested them.

  • The previs is a motion driver, not a render. The camera path, the duration, the timing of the beats and the relationship between subjects all survived. Framing to the pixel and exact geometry did not.
  • Viewport edges hold the set. Same shot, 2 blockings, wireframe visible on one. With the edges the track border stayed continuous to the end; without them the last third drifted. Leave the edges on the geometry the model has to respect, and take them off the ground, where they come back as ruts in the sand.
  • The prompt can take the camera back. The rail held as long as the text did not argue with it. One phrase like as the camera comes closer, or a prop that only reads at a certain size, was enough to trade the framing away.
  • Rehearse cheap, spend once. A frame costs cents, a 720p take a dollar or 2, a 1080p take on the newer model 7. Every decision that can be made on a frame or a cheap take should be made there.

The rejects

These are the rejects: what I threw away along the way, with the reason, because a process is the wrong turns as much as the right ones, and each of these cost something.

47 generations out of 101 were thrown away over the day.

01, the chase: 14 rejected, $29.87

7 frames

7 video takes

02, the drone: 3 rejected, $14.39

1 frame

2 video takes

03, the still life: 15 rejected, $5.54

10 frames

5 video takes

04, the beach: 18 rejected, $10.90

15 frames

3 video takes

The bill

1 day, 4 finished shots, and everything I tried on the way. Here is the bill: what I spent, what came back with a result, and what did not make the cut.

The whole day, 4 shots
$140

$137.38 exactly: 139,595 Magnific credits at $0.00085, plus $18.72 sent to fal.

101paid generations, 45 videos and 56 images
43kept, $71.42 of takes that made the cut
47rejected, $48.81 thrown away, 11 never judged
Where the money went#the-whole-day6 models
Seedance 2.5ByteDance · 11 takes$66.56
Seedance 2.0ByteDance · 34 takes$63.55
GPT 2OpenAI · 16 takes$4.28
Seedream 5 ProByteDance · 17 takes$1.53
Nano Banana ProGoogle · 12 takes$0.76
Nano Banana 2Google · 11 takes$0.70

Note: video is 95% of the bill.

Cost per project#the-whole-day4 shots
$47.08
01The chase
$42.01
02The drone
$28.59
04The beach
$19.69
03The still life
Kept over judged, by model#the-whole-day90 judged
Image models:
GPT 2
7 kept
8 rejected
Seedream 5 Pro
8 kept
7 rejected
Nano Banana 2
4 kept
6 rejected
Nano Banana Pro
1 kept
10 rejected
Video models:
Seedance 2.5
8 kept
2 rejected
Seedance 2.0
15 kept
14 rejected
What one take costs, cheapest to dearest#the-whole-dayas billed
$0.06
StillNano Banana2752 × 1536
$0.11
StillGPT 21344 × 752
$0.95
Video, 4 sSeedance 2.0720p
$1.90
Video, 8 sSeedance 2.0720p
$4.76
Video, 8 sSeedance 2.01080p
$7.34
Video, 8 sSeedance 2.51080p

What comes next

As a 3D artist and as a client, control is essential.

The prompt-only years were frustrating and not much fun: describe a shot, roll the dice, describe it again. This hybrid workflow puts the control back where it belongs. Shapes, lights, forms and a camera move built by hand, then the model dresses it. It is finally fun to play with! And it keeps a 3D artist in the loop for as long as the models need a rail.

What used to take weeks now takes days or hours. Modelling a still life could have been done in an afternoon. Building a miniature city out of the same shapes, a couple of days to a week maybe? Running a full lava simulation on top of it would have been weeks of work between R&D, simulation and rendering. Here it was 1 frame and 1 prompt.

$140 and 1 day bought 4 shots whose camera moves were locked and made with intention, controlled, not hoped for. It's expensive to test, but very empowering at the same time.

Playing with shapes, lights and forms is what made 3D exciting in the first place. That part is still here. But maybe it was never the real fun. Controlling the motion as an animator while the picture takes care of itself is about as good as it gets! A tool set to be aware of, and to play with.

Questions

Does this replace an artist?

No. Not yet. It is a new tool for artists, and eventually for everybody. What it bypasses is the entry, knowing the software, the lighting, the modelling and the control of every detail, which is the time-consuming part. It will still require taste. Some people will keep doing that hard part because they love it, and that's OK. Some will be fine not doing it. The same argument ran when people refused to use a 3D model they had not built themselves, from a marketplace or a library, and today nobody thinks twice about buying a car model and putting it in the shot. This is the same turn. Some will keep the old way, some will try without it, because for the first time in history, there is an option to go right to the result. The quality is not at 100% yet. It is at 80%, and a lot of the time 80% is what the job needs. Good enough, push, deliver, then back to sleep.

Can the model open a scene it did not build?

Untested, and it is the honest next step. Everything on this page was built from an empty document. Opening an existing file and changing it without breaking anything is a different problem.

How long did the building take? About 30 minutes for all the scenes, and most of that was typing. Each scene was about 5 minutes of Claude, Python commands injected one after the other, and it was built. The time went into the camera, by hand, afterwards, and into wrapping my head around what to ask next.

Which model should a shot start on?

Seedream 5 Pro as the default, because it made the most photographic frame that week. That is the honest answer but it changes every week: a new model comes to town and it is the new king, so the answer depends on when the question is asked. What does not change is the freedom of having that many models with that many flavours. Up close they all still have the AI flavour, but with some love I can get closer to my vision. Some are getting close to a photograph, and the closer they get, the sharper the eye that looks at them: the uncanny valley is a long valley to cross. The first Midjourney images from a few years ago read as photographs at the time. Eyes have got better since, and keep pushing that limit. The edge between a photograph and a generated frame gets thinner every month.

Is any of this usable in production?

Yes, with a condition on the client rather than on the tool.

Since last year clients have been building relationships with studios and artists to make content this way, whatever the tool.

Some are firmly against it, some are secretly empowered by it. It is usable in production when the client is happy to test, is fine with the iteration mechanism, and is fine with not chasing every pixel.

Right now that last part is a bit of a roulette, and even so it works for some productions. A few months from now the same models, or the ones after them, will let a small area be isolated and regenerated, and the pixel chasing will come back with it.

What makes it work is the speed. A 3D scene, or an LLM driving a 3D scene, gives a quick draft of geometry, shapes, location, character and motion, which is the shortcut everyone secretly always hoped for.

The final image or video was always the goal; the craft was in the way, knowing the software, having the eye to know what is good. Now the final idea is reachable fast, and the control stays, because the loop is previs, start frame, generate, rinse and repeat. Wrong motion? Fix the previs and resubmit, from the same first frame, it follows. Want another first frame? 2 steps, regenerate the frame from the one that worked, then the video. Iterate as much as needed, knowing the output will follow the motion that was decided and the order of the actions will be respected.

That part is F*^$*-ing fantastic.

Written by Mr. Pinoux, September 2026.