How to Add Image-to-Video, Voice Cloning, and TTS to Rendiv
Rendiv gives developers a useful foundation for building videos with code. React components can define scenes, captions, transitions, and diagrams, while a browser Studio makes the result visible and editable.
A complete production workflow also needs a way to create the media those components display. A product photograph may need to become a realistic moving shot. A script may need narration in the creator’s voice. A revised sentence may require new audio and updated subtitle timing.
Cheap.dev provides APIs for image-to-video generation, voice cloning, and text-to-speech. Combining these services with Rendiv can produce a flexible workflow that supports both human editing and AI agent automation. The key is to connect them through a production layer that manages assets, timing, generation jobs, and validation.
Divide the Responsibilities
A useful architecture has three parts:
| Layer | Responsibility |
|---|---|
| Cheap.dev | Generate speech, clone voices, animate images, and create or enhance media assets |
| Rendiv | Arrange scenes, render captions, animate graphics, and compose the finished video |
| Production layer | Manage generation jobs, store assets, align narration, synchronize timeline edits, and validate exports |
Cheap.dev documents upload workflows, asynchronous generation tasks, status queries, and terminal-state webhooks. These provide the infrastructure for generating assets before they enter the timeline. Cheap.dev API documentation
Keep generation outside React rendering. A component should read a completed media file; seeking through Studio or rendering the same composition again should not submit another paid generation request.
Start with Voice Cloning and TTS
A practical narration workflow begins with a voice sample.
Cheap.dev exposes /api/v1/userVoice/training to create a custom voice from uploaded audio URLs. Once training completes, the resulting voice can be selected through user_voice_id when calling /api/v1/video/send_tts. The TTS response documents an audio URL and generated duration. Cheap.dev OpenAPI
The production sequence should be:
- Upload an authorized voice sample.
- Submit the voice training task.
- Wait for a documented completion state.
- Store the custom voice identifier.
- Generate narration from the script.
- Save the resulting audio locally.
- Measure and align the completed audio before building the final timeline.
Generate narration in natural paragraphs or complete thoughts. This makes revisions cheaper: changing one paragraph should regenerate that paragraph instead of the entire soundtrack.
However, excessively short segments can create inconsistent pacing and intonation. Listen to adjacent sections together, and check that their pauses and delivery feel continuous.
Most importantly, use the actual generated audio duration to schedule scenes. Text length is only an estimate. Different voices, punctuation, and delivery speeds can produce different runtimes from the same script.
Add Image-to-Video Generation
Image-to-video is particularly useful when the available material consists of product photographs, illustrations, or concept images.
A good workflow starts with composition. Prepare the source image for the intended output aspect ratio before animating it. For a landscape video, give the subject enough room in a 16:9 frame and reserve space for captions where necessary.
Then request a short, restrained movement:
A slow camera push toward the product, with subtle natural lighting changes. Preserve the product’s shape, connectors, proportions, and printed lettering.
This is a generation instruction, not a guarantee. Inspect the output for distorted text, changing geometry, disappearing parts, and implausible movement.
Cheap.dev documents image-to-video routes and model-specific controls, including optional end frames and motion-reference modes. Available durations and parameters vary by model, so the integration should validate against the selected endpoint’s schema. Cheap.dev OpenAPI
Generate several short candidates, select the strongest, and trim it in Rendiv. Existing footage should remain the preferred source when it accurately shows the action required.
For movements that merely need a pan or zoom, a deterministic Rendiv animation may be sufficient. Use generative video when the shot benefits from realistic changes in perspective, lighting, or physical motion.
Keep Exact Graphics in Rendiv
Generated media and code-driven graphics complement each other.
Use image-to-video for photographic motion. Use Rendiv for accurate captions, labels, logos, diagrams, and timed arrows. These elements often need exact wording and frame-level timing, which are easier to maintain in editable components.
Adding text after generation also makes localization and revisions simpler. A caption change should not require regenerating the background video.
Build a Reliable Asset Pipeline
Treat each generation as a persistent job.
Record the source asset, prompt, model, parameters, task identifier, status, and output file. Once the result is complete, save it in the project’s asset directory and inspect its dimensions, duration, frame rate, and audio streams.
Cache completed results using the input content and generation settings. Reopening Studio or changing a caption should reuse existing assets.
Webhooks can accelerate completion handling, but they should not be the only mechanism. Cheap.dev describes its terminal-state notifications as best-effort; the task detail endpoint remains the source of truth. Cheap.dev OpenAPI
A timeout after submission also deserves careful handling. Before submitting another paid job, check whether the original task was accepted.
The Missing Piece: Audio Alignment
TTS creates speech, but total audio duration does not provide enough information for accurate subtitles.
A complete workflow needs sentence or word timestamps. Speech recognition can transcribe recorded narration; forced alignment can place a known script against the audio.
The public Cheap.dev contract reviewed for this article did not expose a general-purpose transcription or forced-alignment endpoint. Plan to add a separate alignment service or local tool.
That alignment should drive subtitle timing and help identify useful scene boundaries. When narration changes, rebuild the affected timings instead of leaving the old subtitles attached to new audio.
Preserve Human Timeline Edits
An agent-friendly editor must understand what the human editor changed.
If Studio stores timeline overrides separately from source code, the production layer must read those overrides when planning revisions and exports. Otherwise, an agent may calculate the ending from outdated scene positions.
Use the effective timeline to determine duration:
totalFrames = maximum effective end frame across included tracks
Include intentional pauses and end cards, but exclude unused padding.
Keep time in integer frames wherever possible. Convert audio durations carefully, then verify the exported file: renderer and muxer behavior can still affect the final boundary.
Validate the Export, Not Just the Preview
A browser preview can look correct while the rendered file contains frozen footage, incorrect timing, or missing frames.
Automated checks should cover:
- Unexpected freezes in moving footage.
- Black gaps between scenes.
- Missing assets and fonts.
- Subtitle overflow and unsafe placement.
- Accidental cropping of important subjects.
- Narration cut off at the ending.
- Output dimensions, frame rate, duration, and audio streams.
Also review captions at the size viewers will actually see. A landscape video watched on a phone needs readable text and restrained use of competing titles.
Start with a Small Integration
The first version can be a command-line tool and a scene manifest rather than a new editor.
Implement four operations:
generate_voice: create narration and cache the audio.animate_image: generate candidate video assets.build_timeline: combine asset metadata, alignment, and Studio edits.validate_export: check the rendered output and create review previews.
Add cost estimates before generation. Cheap.dev publishes a machine-readable price catalog at /dev/pricing; match the exact feature and billing unit, and account for multiple candidates and retries. Cheap.dev pricing
With this structure, Rendiv remains the editable composition environment, Cheap.dev supplies generated media, and the production layer keeps the workflow repeatable. The result is a system where both humans and agents can make changes without losing timing, regenerating unchanged assets, or relying on preview alone.
To give Codex or Claude the simplest starting point, share the Cheap.dev skill at https://cheap.dev/SKILL.md and ask the agent to read it before connecting image-to-video, voice cloning, and TTS to your Rendiv project. The skill gives it the API guidance and request examples needed to choose the right endpoints and handle generation jobs.