Claude Opus 5.5 cannot make a video. It takes text and images in and returns text — so the way you make motion graphics with it is to have it write code that a renderer turns into motion. That sounds like a limitation and is mostly an advantage: code is editable, diffable, re-renderable at any resolution, and parameterized, which is exactly what a rendered file is not. This covers which output to ask for, the loop that makes it converge instead of drift, and the one part of motion design it genuinely cannot judge.
Key takeaways
- Claude Opus 5.5 takes text and images in and returns text — it writes the code that renders motion graphics and cannot produce video itself.
- The advantage of code output is parameterization: brand color, duration and copy become variables, so variants and revisions are edits rather than re-exports.
- Released September 22, 2026, with a 1M token context and 128K max output; Anthropic puts it at $4 / $20 per million tokens and claims ~40% lower running cost and ~30% faster output than Opus 5.
- Pick the output format first: SVG and CSS for simple motion, a GSAP timeline for sequencing, a React video framework when you need a real MP4. Do not ask for raw Lottie JSON.
- The loop that converges has a render step in it — generate, render, screenshot specific frames, hand the images back and correct against what actually rendered.
- Adjectives like "punchy" produce the model's average guess. Duration, easing, stagger, anchor point and overlap can all be stated precisely, and they are what decides how motion reads.
- It cannot watch the animation play, so the timing pass is yours. Character animation and a consistent brand motion language are the two things it will not deliver.
- Its knowledge cutoff is June 2026 — paste current library documentation rather than trusting its recall of recent API changes.
What the model actually does here
Opus 5.5 was released on September 22, 2026. Anthropic's platform documentation lists it as accepting text and images and returning text, with a 1M token context window and up to 128K tokens of output. There is no video in or out. Anyone promising otherwise is describing a different product.
What that leaves is substantial. Motion graphics have been expressible as code for years — CSS keyframes, SVG animation, GSAP timelines, Lottie's JSON format, React-based video frameworks that render frames through a headless browser. A model that writes code well can write those, and the thing it produces is a source file rather than an export.
The practical consequence is that your deliverable changes shape. Instead of a finished MP4 you cannot edit, you get a file where the brand color is a variable, the duration is a number, and the copy is a string — so the fifteen-second version, the square crop and the client's revised headline are parameter changes rather than a return trip to the designer.
That is the real case for doing it this way, and it is a narrower case than the enthusiasm suggests. It is strong for templated, repeatable, data-driven pieces. It is weak for anything whose value is in being singular.
What changed with this model specifically
Two things matter for this use, and both are Anthropic's own claims about their own model, so weigh them accordingly.
The first is visual reading. Anthropic describes Opus 5.5 as reading dense documents, charts, screenshots and diagrams at high fidelity. For motion work that is the load-bearing capability, because it means you can render a frame, hand the image back, and ask what is wrong with it. The model is not watching your animation, but it can look at a still from it, which is a different and much better position than describing the problem in words.
The second is cost and speed. Anthropic puts it at $4 per million input tokens and $20 per million output, and says it costs around 40% less to run than Opus 5 while generating output more than 30% faster. Iteration count is what decides whether this workflow is pleasant or miserable, and iteration count is a function of how much a wrong attempt costs you.
One caveat worth holding: Anthropic gives the model a reliable knowledge cutoff of June 2026. Animation libraries move, and it will be confidently out of date on recent API changes. Pin your library versions and paste the current documentation into the conversation when something does not behave as the model expects.
- Text and images in, text out — no video at either end.
- 1M token context, 128K max output, per Anthropic's platform docs.
- $4 / $20 per million input / output tokens.
- Anthropic's figures: ~40% cheaper to run than Opus 5, output ~30% faster.
- Knowledge cutoff June 2026 — paste current docs for anything newer.
If you would rather have this built for you than build it yourself, an AI consultation is where we scope that.
Pick the output format before you prompt anything
This is the decision that determines whether the rest goes well, and most people skip it and accept whatever the model reaches for.
For simple, self-contained motion — a logo assembling, text sliding in, an icon pulsing — ask for SVG with CSS animation. It has no dependencies, runs anywhere a browser runs, and the model is reliably good at it because the output is short and declarative. For sequenced motion where timing between elements matters, ask for a GSAP timeline instead: having one timeline object to scrub and offset is the difference between adjusting a sequence and rewriting it.
For something that must end up as a video file, use a React-based video framework that renders frames through a headless browser. You get real MP4 output, and because each frame is a function of a frame number, the result is deterministic — the same input renders the same output every time, which matters more than it sounds when you are regenerating a piece for the fifth time.
The format to avoid asking for directly is Lottie JSON. It is a fine delivery format and a bad generation target: deeply nested numeric keyframe data where a single misplaced value produces something broken in a way that is hard to read back. If you need Lottie, have the model write code and export to it, rather than writing the JSON by hand.
- Simple, self-contained motion → SVG + CSS keyframes. No dependencies.
- Sequenced, multi-element timing → a GSAP timeline.
- Must be an actual video file → a React video framework rendering real frames.
- 3D or particles → Three.js, accepting that iteration gets much slower.
- Avoid asking for raw Lottie JSON — generate code, export to it.
The loop that makes it converge
Describing animation in a prompt and accepting what comes back produces something plausible and slightly wrong, every time. The loop that works has a rendering step in the middle of it.
Generate the code. Run it. Capture frames at a few points across the duration — the start, the middle of each major move, the end. Hand those images back and say what is wrong with the specific frame. Because the model reads images, the correction is grounded in what actually rendered rather than in your description of what rendered, and that collapses the usual spiral of talking past each other.
Be specific about which frame. "The logo overshoots" is a note about the whole piece; "at frame 18 the logo is past its final position by roughly its own width" is a note the model can act on, and it can see the frame you mean.
Two habits make this much faster. Keep every tunable value — durations, delays, easing, colors, text — in a configuration block at the top of the file, so adjustments are edits to numbers rather than regenerations of logic. And build one element at a time, getting each move right before adding the next; asking for an eight-element sequence in one go produces a piece where everything is slightly off and nothing is isolatable.
Describing motion in terms it can act on
Most prompting advice for this is wrong because it is written as though the model were a designer who will interpret you. It is better treated as a very fast implementer who will do exactly what the brief says, including the parts you left vague.
Adjectives do not survive the trip. "Make it feel premium," "give it some energy," and "make it punchy" produce the model's average guess at those words, which is the same guess everybody else gets. The things that actually determine how motion reads are duration, easing, stagger, anchor point, and overlap — and all five can be stated precisely.
So instead of "snappy entrance," write: 320ms, ease-out with a slight overshoot, scaling from 0.92 to 1.0, anchored at the center, each of the five items starting 60ms after the one before it. That is a brief with one interpretation. If you do not know those numbers yet, ask the model to produce three variants at different timings and look at them — picking from options is a faster way to find the feel than describing it.
The one adjective worth keeping is a named reference, because it carries real information the model has seen plenty of: "the way a system notification slides in on macOS" specifies a curve and a duration far more tightly than "smooth" does.
- Duration in milliseconds, not "quick".
- Easing by name or curve, not "smooth".
- Stagger as a per-item offset, not "one after another".
- Anchor point explicitly — most wrong-looking scale animations are anchor bugs.
- Overlap: say whether the next move starts before the previous one finishes.
Where it genuinely falls down
The hard limit is that it cannot watch the animation play. It sees still frames you give it, and a great deal of what separates good motion from adequate motion exists only in time — whether a move lands or floats, whether two elements are fighting, whether the whole thing has a rhythm. A sequence of correct frames can still play badly, and the model has no way to know that.
This is why the final pass is yours and why it is not a formality. Expect to adjust timing by eye after the structure is right, and expect that to be the part no amount of better prompting removes.
Character animation is the second wall. Anything with weight, anticipation and follow-through — a figure moving in a way that reads as alive — is not a code-generation problem, and attempting it this way produces the uncanny stiffness you have seen in a hundred auto-generated explainers.
The third is a consistent motion language across a body of work. A brand's motion identity is a set of decisions applied the same way every time, and a model starting fresh each session will not hold them. The fix is not better prompting; it is writing the decisions down — your durations, your easing curves, your stagger — and pasting them in every time, which is a style guide doing what style guides have always done.
When to hire a motion designer instead
A few cases where this is the wrong tool, and recognizing them early is cheaper than discovering them late.
If the piece is the product — a launch film, a brand identity, anything whose job is to be memorable — hire someone. The value there is in judgment and singularity, which is precisely what a generated piece lacks. You will spend longer fighting toward distinctive than a professional would spend making it.
If you need one piece, once, hire someone or use a template. The payoff of the code approach comes from parameterization and repetition — forty variants, data-driven updates, weekly output. For a single asset you are paying a setup cost with nothing to amortize it against.
And if nobody on your side can read the code, be careful. A generated animation that breaks in Safari, or needs a change in eight months, is a liability when the only person who understood it was a chat session that has since been closed. That is the same trap as any generated code, and it applies here with the added difficulty that visual bugs are harder to describe than functional ones.
- The piece is the product → hire a motion designer.
- One asset, once → a template is cheaper than a setup cost you cannot amortize.
- Nobody can read the output → you are accumulating something you cannot maintain.
- It needs to feel alive → character motion is not a code-generation problem.
A first piece worth building
If you want to find out whether this fits your work, do not start with the hero animation. Start with the thing you need forty of.
A good first build is a templated social card that animates: your logo, a headline, a statistic, a background treatment, in your brand colors, where every one of those is a variable at the top of the file. Get one right, then generate the next twenty by changing strings and numbers.
That exercise answers the real question quickly. If the twenty variants are genuinely useful and took minutes each, the approach fits your operation and it is worth building the template properly. If you spent the whole time fighting the first one toward something you would actually publish, that is your answer too, and it cost you an afternoon rather than a project.
Either way, keep the file and the configuration block. The asset is disposable; the template is the thing that was worth making.
The same caution applies here as to any code you did not read — where accepting generated output stops being safe covers the line between a prototype and something people depend on.
Related services
Web Design
We design and build high-performance websites for AI-era businesses — fast, semantic, and engineered to turn first-time visitors into paying clients.
Explore premium webSocial Media Automation
We build content systems that draft, schedule, publish, and route replies across your channels — so the calendar stays full whether or not anyone remembers to post.
Explore content engine
