Every AI lip-sync model on the same picture and the same song, side by side

This page sends one still picture and one piece of a song to every audio-to-video model in the Creator Studio, the models that make a picture talk or sing: Wan 2.2 S2V, OmniHuman 1.5, HeyGen Avatar IV, Kling Avatar v2, VEED Fabric 1.0 and SadTalker. Each clip is shown with its token cost, its length and how long it took to make, so you can compare lip-sync quality and price on exactly the same input.

The tests come from a real music video made on RFab, Don't Toy with Me, Miss Malfoid: 2D-anime stills and the finished song, sung, with the band under it. Two are the ordinary job, one singer close up and one singer waist up with a big pose. Two more redraw the first still as a photograph of a real person and as a western TV cartoon, with the same pose and the same song, so you can see how each model handles realistic and cartoon faces against anime. Three put more than one face in the frame: a second girl walking past behind the singer, two girls side by side where only one shouts, and four girls with a backup echo. Multi-face shots are where these models fail most, and they fail in different ways, so each test says what to watch for.

To try your own, pick the models, use a test's picture and audio or upload your own portrait and an audio clip (5 to 60 seconds), and run them together. Lip-sync is billed per second of audio with a 5 second minimum, and the page shows the estimated total before you confirm. Every clip lands in your gallery.

Open AI Lip Sync Comparison on Reality Fabricator

What you can do

How it works

  1. Pick a test: one singer, or a shot with several faces in frame.
  2. Play the clips and compare mouth shapes, timing, body motion and whether the right face sings.
  3. Tick the models you like and open the try-it-yourself tab.
  4. Use the test's picture and audio or upload your own, confirm the estimate, and compare the clips.

Frequently asked questions

Which AI lip-sync model is best for singing?

After watching every model on these tests, the owner ranked them OmniHuman 1.5, Wan 2.2 S2V, VEED Fabric 1.0, Kling Avatar v2, HeyGen Avatar IV, then SadTalker. Wan is almost as good as OmniHuman at about a sixth of the price, but adds erratic movements that can look unnatural. When a second character is in the shot and must stay quiet, Wan is the only model that keeps them from mouthing the words.

Can lip-sync models handle more than one person in the picture?

Only Wan 2.2 S2V reliably keeps a second character quiet while the singer sings. The others are built for one face: with several faces in frame some move every mouth, some change the other character, and SadTalker crops to a single face, often the wrong one. With three or more faces even Wan lets more than one mouth move, so for a group, cut between single-singer shots.

How is lip-sync priced?

Per second of audio with a 5 second minimum, at each model's own rate shown on its card, from about 1,650 tokens a second for Wan 2.2 S2V to about 11,500 for VEED Fabric. SadTalker is one flat price per clip. Viewing the demo costs nothing.

Do lip-sync models work on anime, cartoon and realistic pictures?

Yes. Most tests here use 2D-anime stills, and the first one is repeated as a photograph and as a western TV cartoon with the same pose and song. Most models keep the character on model; SadTalker works but crops to a small square around one face, and HeyGen Avatar IV tends to make odd mouth shapes on anime faces.

What do I need to try my own?

An account with tokens, one picture with a clear face, and an audio clip between 5 and 60 seconds of speech or singing. Each clip is billed at the listed rate and saved to your gallery.

Related tools

Which lip-sync AI to use

Singing, best to worst, as ranked by RFab's owner watching the clips with the song: 1. OmniHuman 1.5 (10,770 tokens a second), 2. Wan 2.2 S2V (1,650 tokens a second), 3. VEED Fabric 1.0 (11,540 tokens a second), 4. Kling Avatar v2 (4,310 tokens a second), 5. HeyGen Avatar IV (2,970 tokens a second), 6. SadTalker (38,000 tokens a clip). Wan 2.2 S2V is almost as good as OmniHuman at a sixth of the price but adds erratic movements that can look unnatural, and it is the only one that keeps a second character in the shot from talking.

Default: OmniHuman 1.5 for a sing shot with one character in frame (the best). If anyone else is on screen and must not talk, ALWAYS Wan 2.2 S2V: it is the only model that keeps them quiet. Wan is also ALMOST as good as OmniHuman at a sixth of the price, but it adds erratic movements that can look unnatural: the budget pick, or the retry when OmniHuman gets a shot wrong. Then VEED Fabric, Kling Avatar v2, HeyGen Avatar IV (never on anime), SadTalker last. For movement or several people, do not lip-sync at all: make a DANCE shot.

Multi-face test (rfab.ai/compare/lip-sync, Oct 11 2026): with two or four faces in frame most models also moved a non-singer's mouth; Kling Avatar v2 walked a background girl out of frame and brought in a different one; SadTalker cropped to the wrong face every time. Only Wan 2.2 S2V keeps a second character quiet (Collins, Oct 11 2026): use it whenever someone else is in frame and must not talk.

OmniHuman 1.5 (the default singModel)

Best at: SINGING and performance: ranked BEST by Collins (Oct 11 2026). Gestures, head tilts and shoulders that follow the song, with the clearest vowel shapes. The default singModel.. Use it when: every sing shot with ONE character in frame by default, and any shot where the body should move with the music. Not when: ANOTHER CHARACTER IS IN FRAME AND MUST NOT TALK: use Wan 2.2 S2V (Collins, Oct 11 2026: only Wan keeps the second character quiet). Also not when the hands hold something that must stay exactly put, or the budget cannot take it (6.5x Wan 2.2 S2V).

Limits: 5-60 s of audio (best under 15 s); one performer; it moves the arms and head a lot, so a still with props in the hands can drift; 1472x832 out. Speed: about 2 to 3.5 min for 5-10 s.

Wan 2.2 S2V

Best at: ALMOST as good as OmniHuman 1.5 at a sixth of the price, and the ONLY model that keeps a second on-screen character from talking. Ranked SECOND by Collins (Oct 11 2026).. Use it when: ALWAYS when a second character is on screen and must not talk: Wan is the ONLY model that keeps the other character quiet (Collins, Oct 11 2026). Also the budget pick for sing shots (a sixth of OmniHuman's price), or a shot OmniHuman got wrong. Not when: a still whose hands must stay exactly as drawn (they can smear).

Limits: 5-60 s of audio; one face; a prompt is always sent (it drives the motion); HANDS: if the still shows hands it animates them and they smear or go see-through (film 51acae1d, shot i_never, Oct 10 2026, Collins: "she loses her hands ... they go see-through"). For a Wan sing shot, frame head and shoulders with the hands out of the picture, or hands resting still and say so in the motion.. Speed: about 1.2 to 2.5 min for 5-10 s.

VEED Fabric 1.0

Best at: a steady talking or singing face that keeps the still almost exactly as drawn; ranked THIRD by Collins (Oct 11 2026).. Use it when: the still must barely change and OmniHuman and Wan both moved it too much. Not when: a performance with body motion (OmniHuman); or on price, it is the dearest per second.

Limits: 5-60 s of audio; one face; 720p (1312x736); little body motion. Speed: about 1 to 1.5 min for 5-10 s (fast).

Kling Avatar v2

Best at: a SING close-up of one face that keeps the pose of the still; ranked FOURTH by Collins (Oct 11 2026).. Use it when: the three above have failed on a shot. Not when: a shot with anyone else in frame: on the two-face test it walked the background girl out of the shot and later brought in a different one.

Limits: 5-60 s of audio; one face (every face it finds may move its mouth); keeps the pose of the still, little body motion; 720p out; the slowest. Speed: about 4 to 5 min for 5-10 s (the slowest).

HeyGen Avatar IV

Best at: a cheap presenter-style talking head of a REALISTIC person; ranked FIFTH by Collins (Oct 11 2026).. Use it when: dialogue or narration to camera by a realistic person, where the pose must not change. Not when: ANIME / 2D CHARACTERS: do not use it. On an anime character it makes weird mouths (Collins, Oct 10 2026, film 51acae1d shot i_never: "heygen for anime it makes weird mouths"). Also not for singing with big vowel shapes, or for several faces in frame (on the four-face test every girl sang).

Limits: 5-60 s of audio; one face; 720p out; upper-body motion is small (the hands drift a little). Speed: about 1 min for 5-10 s (fast).

SadTalker

Best at: nothing a music video needs: ranked LAST by Collins (Oct 11 2026).. Use it when: never for a film. Only a quick, low-resolution talking face where the framing does not matter. Not when: any shot that must keep its framing, and any picture with more than one face: it cropped to the WRONG face in all three multi-face tests (the non-singer, or a face in the crowd).

Limits: crops to a small square around ONE face (256 px out): the rest of the picture is gone; one flat price per clip whatever the length. Speed: about 20-70 s (the fastest).

The 6 models compared

ModelProviderCost per second of audioBest for
Kling Avatar V2 — Talking Headreplicate4,310 tokens ($0.086), 5 s minimumTurn any portrait into a realistic talking avatar with audio-synced lip sync and natural expressions. Billed per second of audio (min 5s, max 60s).
SadTalker — Talking Head (realistic styles only)replicate38,000 tokens ($0.76) per clipSingle portrait image + audio = talking head video. Needs a clear, REALISTIC human face — its face detector fails on anime/stylized portraits (use Kling Avatar for those).
Wan 2.2 S2V — Audio-to-Videoreplicate1,650 tokens ($0.033), 5 s minimumConvert audio clips and reference images into synced talking head videos with natural motion.
OmniHuman 1.5 — Talking Head (full body, gestures)replicate10,770 tokens ($0.22), 5 s minimumByteDance OmniHuman 1.5: one portrait or full-body image + audio → a talking video with semantic gestures and body motion. Best under 15s of audio (quality degrades past that). Billed per second of audio ($0.14/sec).
VEED Fabric 1.0 — Talking Head (720p)replicate11,540 tokens ($0.23), 5 s minimumVEED Fabric 1.0: any image + audio → a 720p talking video. Handles stylized portraits. Billed per second of audio ($0.15/sec at 720p).
HeyGen Avatar IV — Talking Headrunware2,970 tokens ($0.059), 5 s minimumHeyGen Avatar IV: one photo + audio → a talking (or singing) video with expressive face and upper-body motion. Billed per second of audio ($0.0385/sec at the provider; min 5s, max 60s).

The tests and what each model made

One singer, hands up

The still every model was given: One singer, hands up The audio (5 s): I had no idea it was only a bluff, but grandeur went white. One face.

Motion line sent with it: She sings with an innocent shrug, then glances sideways and bites her lip.

What to look for: One anime face, medium close-up, both palms up. Watch the mouth on "only" (a round "o") and the breath before "but" (the mouth should close), and whether the raised hands stay hands.

6 models made it; the fastest was SadTalker — Talking Head (realistic styles only) (25 s).

Photoreal: the same singer as a real person

The still every model was given: Photoreal: the same singer as a real person The audio (5 s): I had no idea it was only a bluff, but grandeur went white (the same 5 seconds as "One singer, hands up"). One face.

Motion line sent with it: She sings with an innocent shrug, then glances sideways and bites her lip.

What to look for: The first test redrawn as a photograph: same pose, same song. Real teeth, lips and skin are where a fake mouth shows first. Does the face stay the same woman, and do the raised hands stay real hands? Compare each model with its anime card.

6 models made it; the fastest was VEED Fabric 1.0 — Talking Head (720p) (48 s).

Western cartoon: the same singer, American TV style

The still every model was given: Western cartoon: the same singer, American TV style The audio (5 s): I had no idea it was only a bluff, but grandeur went white (the same 5 seconds as "One singer, hands up"). One face.

Motion line sent with it: She sings with an innocent shrug, then glances sideways and bites her lip.

What to look for: The first test redrawn as a western 2D cartoon: thick outlines, flat colour. Does the mouth move like a cartoon mouth, or does a realistic mouth get pasted onto a flat face? Do the outlines stay clean and the style stay flat instead of drifting to 3D?

6 models made it; the fastest was SadTalker — Talking Head (realistic styles only) (16 s).

One singer, waist up, a big pose

The still every model was given: One singer, waist up, a big pose The audio (9.9 s): "She's never been wrong, just ask, she'll say" (with the line before it). One face.

Motion line sent with it: She sings to the camera full of sass: flicks her hair, sways her hips, tilts her head with a smug grin.

What to look for: A waist-up performance: does the body move with the song or only the mouth, does the hand in her hair stay a hand, and does the face stay the same girl for ten seconds.

6 models made it; the fastest was SadTalker — Talking Head (realistic styles only) (24 s).

Singer with a second girl behind her

The still every model was given: Singer with a second girl behind her The audio (9.9 s): the same ten seconds as "One singer, waist up". 2 faces in frame, one voice.

Motion line sent with it: She sings to the camera full of sass: flicks her hair, sways her hips, tilts her head with a smug grin. The girl behind her does not sing.

What to look for: The same still and song as the waist-up test with ONE change: Granger walks past behind her, out of focus. Does only the singer sing, or does Granger start mouthing the words too? Compare each card with its waist-up twin.

What we saw: Kling Avatar v2 walked Granger out of the shot and later brought in a different red-haired girl; OmniHuman 1.5 had someone walk through the frame at the end; SadTalker cropped to Granger's face instead of the singer's.

6 models made it; the fastest was SadTalker — Talking Head (realistic styles only) (16 s).

Four faces, one lead voice

The still every model was given: Four faces, one lead voice The audio (6.5 s): "Oh, Granger, (Granger!) I'm just curious, why are you so cute when you're furious?". 4 faces in frame, one voice.

Motion line sent with it: The blonde girl in the middle sings into the eclair; the two girls beside her echo "Granger!"; the girl below looks up, furious.

What to look for: Four faces and a chorus with a backup echo. Which mouths move: the lead only, everyone, or the wrong one? Does anybody sing with an eclair in her mouth? A multi-face shot is where these models fail most, and how they fail differs.

What we saw: Kling Avatar v2 moved the lead's mouth most and the others least. Wan 2.2 S2V and HeyGen Avatar IV had the girls beside her singing too; OmniHuman 1.5 moved the brunette on the left more than the lead; VEED Fabric barely moved any mouth; SadTalker cropped to a face in the crowd behind them.

6 models made it; the fastest was VEED Fabric 1.0 — Talking Head (720p) (64 s).

Two girls side by side, one shouting

The still every model was given: Two girls side by side, one shouting The audio (5 s): "Get out. GET OUT." then the chorus "Oh, Granger, (Granger!)". 2 faces in frame, one voice.

Motion line sent with it: The blonde girl in pyjamas shouts and points down the corridor; the girl with the tea tray smiles calmly and says nothing.

What to look for: Two clear, equally sized faces. The blonde shouts; the girl with the tray should stay quiet. Does the model pick the right face, both, or neither, and does the pointing arm survive?

What we saw: Most models opened the quiet girl's mouth as well; HeyGen Avatar IV kept it the stillest. SadTalker cropped to her face, the wrong one.

6 models made it; the fastest was SadTalker — Talking Head (realistic styles only) (24 s).

Back to Reality Fabricator