Video annotation

From footage to training data

Import supported CFR or VFR video into any project and use the workflow that matches its task: box tracks for object detection, object-mask tracks for instance segmentation, frame-by-frame keypoints, or frame ranges for classification.

The AnnotateIt media grid showing dataset thumbnails with annotation status indicators.

Related video tutorial

See this workflow in action

Turn Video Into a Dataset

Extract frames, drop duplicates, split, export

Video without the upload step

  • Keyframe object tracks

    Object Detection keeps box identity across keyframes; Instance Segmentation keeps an object-mask track. Boxes and compatible polygons interpolate; raster masks hold their exact pixels until the next keyframe.

  • Local frame extraction

    Frames are selected using timing read from the video, not an assumed 30 FPS, and decoded on your device. Nothing is sent anywhere.

  • Annotate like images

    Keypoints stay frame-by-frame; the three Classification variants label whole-frame ranges.

  • Annotations come along

    Extracting frames copies the annotations you already drew — interpolated track positions included — onto the resulting images, by default.

  • Take photos or record video

    Take photos manually, on a timer, after motion or when Smart capture finds a useful frame — or switch to Video and record a silent clip from the selected camera. Photos and completed clips stay staged on-device until you review them and select Accept.

  • Progress at a glance

    The media grid shows the annotation status of every frame, so you always know what is left.

  • Video-aware export

    Keep original video bytes, frame annotations and tracks in video-aware formats such as Supervisely Video, MOT, MOTS or Datumaro. Extract frames first for image-oriented COCO, YOLO or Pascal VOC workflows.

All six image/video project types accept images and supported CFR or VFR video, including compatible Camera clips. Object Detection uses keyframe box tracks, Instance Segmentation uses object-mask tracks, Keypoint Detection is frame-by-frame, and the three Classification variants use frame ranges. You can also extract selected frames as images locally.

Choose how footage enters the workflow

A recording can become a native video dataset or a source of selected stills. Choose according to what you need to label.

Import the video

Timeline annotation

Keep the clip for tracks, frame annotations and classification ranges. Preview sampled model predictions or extract frames with their annotations.

Play it as a capture source

Select useful still images

Use playback speed, looping and manual, timelapse, motion or available Smart capture rules. Inspect staged frames before Accept; their recording positions stay with the images.

Collect new footage

Camera or supported screen sharing

Collect stills or record silent clips. Source names, sessions and capture reasons connect the resulting media back to its collection.

RTSP/HTTP network cameras are available in the Windows and macOS app; watched folders are not implemented. A playable local recording can be used as a video-file source.

Common questions

Can I select frames while a local video plays?

Yes. Open a playable video file as a Capture source and collect stills manually or with timelapse, motion or available Smart rules. Captures stay staged until Accept and retain their recording timestamp. This differs from importing the video for timeline annotation.

Can a model preview predictions across a video?

With an available model in AI mode, precompute sampled predictions for playback. Inspect them and explicitly accept the current frame or cached predictions across the video. This is not frame-locked real-time inference or automatic object-identity tracking.

Which video files can I open?

MP4, MOV, M4V and WebM with CFR or VFR, readable timing and a codec supported by the running platform. AVI/MKV frame extraction is available through the native desktop decoder, but playback can still be unavailable. Variable-frame-rate video is supported with average FPS and exact frame indices; imported bytes and timing stay unchanged.

Is the video uploaded for processing?

No. Frame extraction happens on your device. Sending a dataset to an optional ML runner is a separate, explicit action.

Do I have to annotate every frame?

No. You pick the frames worth labeling — nothing forces you through the whole clip.

Can I track one object across frames?

Yes, in Object Detection and Instance Segmentation. Detection keeps box identity across keyframes; instance segmentation uses object-mask tracks. Keypoint Detection is frame-by-frame instead, while the three Classification variants label frame ranges. There is no automatic tracker model; keyframes are yours to place.

Do I lose my work when I extract frames as images?

No. Extraction copies the annotations of any frame that has them onto the extracted image — it is on by default, and it happens before the video is removed if you asked for that.

Which formats can a video dataset export?

Supervisely Video, MOT and MOTS apply when the dataset contains video — MOT for detection tracks, MOTS for instance-segmentation masks — while Datumaro can mix images and native video. For image-oriented COCO, YOLO or Pascal VOC, extract frames first. Untouched original video bytes stay unchanged in exports.

Why is an older annotated video blocked?

Older videos without verified frame timing are blocked to keep labels from moving onto different frames. Their ids and annotations are unchanged. Re-import the original as a new video and review any annotation transfer separately; an old annotated archive does not silently upgrade the timeline.

Open AnnotateIt in your browser

Create your first local project in the browser. Optional ChatGPT or Claude annotation needs your own AI connection.

Questions before you start?Contact support →

Video tutorial

AnnotateIt tutorial