Video & AI

Tracking & footage

Region motion tracking, a 3D camera solve, and everything footage — trims, speed ramps, unlinked audio, speech and auto-captions.

Motion tracking

Tracks in Kreativus are things, not actions: stored on the footage layer, saved with the project, re-solvable and re-applicable — to more than one layer, more than once. The Tracking panel is where they all live, grouped by clip.

Two kinds, two promises. A Point answers “where is that” — click one spot (a freckle, a bolt, a window corner) and it follows position only, because over something that small, rotation and scale are noise. A Surface answers “where is it and which way is it facing” — click the four corners of a flat thing (a sign, a phone screen, a box) and it follows movement, turn and growth, and can carry another layer into the shot’s own perspective.

Under the hood it’s a region tracker: dozens of features scattered in your box, a consensus fit, outliers voted away — so an occluded corner or a passing shadow costs one point, not the track. Every feature is verified against its original appearance, never frame-chained, which is what stops the classic slow slide off the target. Confidence sets how sure a frame must be (every finished run reports its worst frame — set the bar under that to catch it), and When unsure picks the honest failure: Stop there (keep only what it actually saw) or Keep going, holding position. Brief dropouts — a blur, something crossing — are bridged as a smooth run between the two frames it did measure.

◀ / ▶ solve from the pins, and frames merge into what’s already solved — so the fix loop is: scrub to where it slipped, drag the pins back on, carry on from there. Then apply, as many times as you like:

  • Follow → Control layer — a null with a visible marker carrying the motion (and, for surfaces, rotation and size — each a per-child Inherit checkbox). Parent anything to it.
  • Stick (surfaces) — corner-pins the selected layer onto the tracked surface, its own artwork landing on the pins every frame. A sign, a screen, a graphic sitting in the shot — and the pins stay draggable afterwards if a frame wants a nudge.

Camera solve

Region tracking answers “where did that thing go on screen.” The 3D camera solve answers where the camera was, in space, on every frame — so an extruded title or a 3D group can be placed into the shot and stay put while the camera moves around it.

Press This shot. It reads the whole frame (no pins), watches the feature cloud build as it works, and produces: a real Camera layer with keyframed position, rotation and solved focal length; the scene’s sparse point cloud (kept — it’s what Depth Mask anchors to); and, on request, Markers — nulls dropped at solved points spread through the scene, to parent things to. The result line tells you everything worth knowing before trusting it: how many frames placed, the lens (and any barrel/pincushion it measured), how far features land from where they were seen, and whether it found a real floor or took level from the camera.

One physical requirement, stated plainly when it bites: the camera must move through space. A pan or a zoom from one spot has no parallax — nothing in the shot has depth to find — and no tracking quality can fix that.

Footage

Import .mov, .mp4, .m4v (image sequences too). A clip’s in-point trims without disturbing keyframes, and Time Remap is the speed system: no keys = native speed, correctly rate-converted to the comp; keyframe it and you have speed ramps, freeze frames, and reverses — the value is the source frame.

Audio rides along: waveform on the layer bar, a draggable volume curve directly over it (keyframable), dB nudges, fade in/out buttons, per-layer audio effects, mute/solo — and Unlink Audio from Video extracts the clip’s track onto its own independently-trimmable layer. During scrubbing the decoder drops to draft resolution automatically, which is why heavy footage stays interactive.

Audio & captions

The Audio layer speaks, records or plays: Text to speech types a line and has a system voice read it (editable and re-generatable later), Record performs against picture (a 3-2-1 count-in, the comp playing back from the playhead — Premiere-style), Instrument gives you a playable music track. Audio files import like any other file — drop them in or use File → Import.

Generate Captions auto-transcribes any audio layer or video’s audio — on-device, offline, free — into a real Captions layer, timed to the word. The Captions panel is a proper transcript: every cue an editable row, click a time to jump there, click between word chips to split a cue exactly at that word. Caption styling runs deep — per-word karaoke highlight colour, arrival animations, shadows, backgrounds, and a preset gallery.

Recipe — a label that lives on a moving box. Surface-track the box’s face (four clicks) → ▶ → select your label layer → Stick. Scrub it; if one frame wobbles, nudge the pins there and re-solve — the fix merges in.

The Footage workspace

The workspace switcher has a Footage preset for video work: the Tracking panel holds a permanent slot above Properties (a track is applied to a different layer than it was tracked on, so it earns its own always-visible home), with the Timeline below for scrubbing. Its scope leads with the video, animation and simulation tools.

Wrong, unclear, or missing? Tell us — docs bugs count as bugs.