You want one cinematic shot, with sound. Use Veo. Its output is generated shots with audio, and for a single shot where fidelity matters more than structure, that is exactly the right purchase.
You want to explore an idea visually, fast. Use Sora. Prompting from text or an image inside OpenAI's app is the lightest way to see what an idea looks like before committing any production to it.
You want a finished episode instead of a shot. Neither model does that, by design. That is where an episode pipeline like Frutti fits: one written premise becomes a scripted, voiced, captioned 9:16 episode with a cast that returns, while the models remain the better buy for raw footage.