sync.so just launched something that changes how video gets made. Their sync-3 model can now take a single still image, any still image, and turn it into a fully lip-synced talking video. No source video needed. One photo, one audio track, and out comes a person speaking with accurate mouth movement, natural head motion, and facial expression that matches the words.

This is not a separate tool bolted onto their existing product. It is the same sync-3 model that already handles video lip sync, now asked to start from less. Where the old workflow required existing footage of someone talking, image-to-video builds the entire performance from one frame.

What Image-to-Video Actually Does

Upload a JPEG, PNG, or WebP of a person. Add an audio track, whether that is a cloned voice, a generated voice, or a recording you already have. sync-3 analyses the face in the image, builds a global understanding of the person across the whole shot, and generates every frame at once. The output length matches your audio. A 30-second script gives you 30 seconds of video.

The model works across 95+ languages. A single portrait can deliver a line in English, then the same line in Hindi, then Japanese, off one image. If there is more than one face in the frame, you tap the one you want to speak.

What makes this different from every other image-to-video tool: sync-3 does not work in small independent chunks around the mouth. It reasons about the whole face, what to move and what to hold, generating every frame holistically instead of patching regions together. That is why it can start from one frame. The model has enough understanding of facial structure to construct realistic motion from the smallest possible input.

The Creative Unlock

For anyone making content from generated stills, this collapses a whole production step. You used to need an image model, then a separate motion or video model, then a lip sync pass. Three tools, three exports, three rounds of quality loss. Now the image is the only thing you have to make first.

A Midjourney character can speak. A headshot becomes a presenter. An illustration delivers a script. A painting talks. Any visual that shows a face can become a speaking video.

For small businesses making product videos, social content, or marketing explainers, this means you can create a presenter video without ever picking up a camera. Generate or commission a portrait, write a script, add a voice, and you have a talking-head video ready to post.

The DaVinci Resolve Plugin

The same launch added a native DaVinci Resolve Studio plugin. This is the second NLE plugin from sync.so, following their Premiere Pro integration earlier this year.

What it does: select a clip on your Resolve timeline, hit generate inside the plugin, and the lip-synced version lands back in your edit. No export. No browser tab. No round trip to a web app. You pick in and out points, generate, preview in place, and drop the result into your timeline when it looks right.

The plugin runs every sync.so lip sync model, including sync-3, which reads the whole scene instead of a cropped region around the mouth. That matters for finishing sessions: sharp angles, low light, multiple speakers in frame, close-ups at up to 4K 60fps. The performance in the original take carries through instead of getting flattened.

What This Unlocks for Content Creators

Three workflows just got shorter:

  • Video localisation now starts from one image instead of one video per language. Generate your presenter image once, add translated audio for each market, and produce localised talking videos without reshoots.
  • Social content from stills is now a single step. Your product shot, your brand portrait, your AI-generated character can all become short speaking videos without shooting anything.
  • Post-production fixes inside your NLE are now possible. Need to fix a line of dialogue in Resolve? Select the clip, generate the corrected lip sync, and drop it back in. No export-import cycle.

How to Try It

Image-to-video is live now for everyone on sync.so. Open the studio, upload an image where you would normally upload a video, add your audio, and generate. That is the whole flow.

The DaVinci Resolve Studio plugin is a free download. Install it, and run your first sync from inside Resolve Studio.

For best results with image-to-video, use a clear, reasonably front-facing image of a person. The model does the rest.

Want more like this? Join MOKU Club for free. Weekly resources, early access to new guides, and occasional templates you can actually use. Join below.

Person typing on laptop in dark minimalist studio, representing creative AI workflows
Adobe Firefly Graph Just Turned Creative Workflows Into Reusable Assets. Here Is What That Actually Means.AiDesignNews

Adobe Firefly Graph Just Turned Creative Workflows Into Reusable Assets. Here Is What That Actually Means.

HeraHeraJune 18, 2026
Vidu S1 Just Turned a Single Photo Into a Live Video Call With an AI CharacterAiDesignNews

Vidu S1 Just Turned a Single Photo Into a Live Video Call With an AI Character

HeraHeraJuly 7, 2026
How to Build a Complete Product Page in Under 20 Minutes With Shopify TinkerDesignHow-toSmall Businesses

How to Build a Complete Product Page in Under 20 Minutes With Shopify Tinker

HeraHeraApril 29, 2026