No, the video is not uploaded when you extract a video frame using a browser-based tool like Video Frame Extractor. The source clip stays on your device from start to finish: the browser opens the local file, the media element decodes the audio and video tracks, the decoder seeks to the requested time, and the resulting still is drawn to a canvas, encoded as a PNG, and saved through a local download link. There is no upload pipeline, no cloud transcoder, and no server-side render at any point in that path. The same applies whether the file is a screen recording, a phone clip, or a downloaded lecture — every byte is read from disk and processed by the browser's own decoder before any image is produced. This article explains exactly where the video goes during extraction, why the result still feels fast even though no file is sent anywhere, and the practical limits — file size, duration, dimensions, codec support — that determine whether the browser can complete the extraction entirely on your device. If any of those limits are not met, the tool reports a decode, seek, canvas, or encoding failure instead of producing a silently empty PNG.

Where the Video Goes During Frame Extraction
The simplest way to answer the upload question is to follow the file. When you pick a file in an extractor, the browser hands the page a File object created from a local Blob. That object is fed into a <video> element through a temporary object URL. From that moment on, every operation — duration, current time, decoded frame — runs inside the browser's media stack on your tab. No network request is made for the video bytes themselves.
You can confirm this in practice by opening the browser's network panel before and during extraction. Only the page itself and its small static assets load; the chosen video never appears in the request list. The reason a server does not feature in the pipeline is that there is no reason to send the file away. A PNG is a single image, and the browser already holds the decoded frame in memory the moment the media element finishes seeking. The tool only needs to copy those decoded pixels onto a canvas, encode the canvas as a PNG, and expose the result through a revocable URL.createObjectURL download. That whole chain runs on the CPU and GPU of your own device.
For readers who want a deeper walkthrough of the privacy side, the guide Can I Extract a Video Frame Without Uploading a File? covers the same boundary from a different angle.
How Local Extraction Works Inside the Browser
Frame extraction in a browser is built from four standard web platform primitives: an HTMLMediaElement for decoding, a currentTime property for seeking, a CanvasRenderingContext2D.drawImage call for copying pixels, and a PNG encoder — typically canvas.toBlob — for output. The MDN reference for HTMLMediaElement.currentTime describes how setting that property moves the playhead to a requested media timeline position, and the MDN reference for Canvas drawImage explains how the current decoded frame is rasterized onto the canvas at its native dimensions.
Putting those primitives together produces a clean extraction path that the tool implements one step at a time:
- The video element loads a local File through an object URL, with no network transfer.
- The tool reads the element's duration to validate the requested time against the reported timeline.
- The tool assigns video.currentTime = t and waits for the seeked event before reading pixels.
- A same-size canvas is created and the current frame is drawn onto it with drawImage(video, 0, 0).
- The canvas is encoded as a PNG and offered as a local download through an object URL.
Stale object URLs are revoked when inputs change or the component unmounts, so the underlying file buffer does not linger in memory beyond its useful life. Decode, seek, canvas, and encoding failures are reported explicitly instead of producing an empty download.
File and Codec Limits That Affect Extraction
Because extraction depends entirely on what the browser can decode, the operating limits are the ones the browser can actually handle. Video Frame Extractor applies a shared video safety policy that constrains the input as follows:
| Limit | Value |
|---|---|
| Maximum file size | 500 MiB |
| Maximum duration | 5 minutes |
| Maximum dimension per side | 4096 pixels |
| Maximum pixel area | 3840 × 2160 = 8,294,400 pixels |
| Supported containers | MP4, WebM, MOV, M4V, Ogg |
The area cap is worth pausing on. Multiplying the two sides directly gives 3840 × 2160 = 8,294,400 pixels, which is roughly 8.29 megapixels. That is enough for any single still you would realistically grab from a short clip, but it rules out working with long, full-resolution cinema footage inside the browser.
Browser codec support is the other hard limit. A file extension such as .mp4 or .mov only describes the container; the audio and video inside that container are encoded with specific codecs — for example H.264, HEVC, VP9, or AV1. The browser has to support those codecs in order to decode the file at all. A file that opens in one browser may fail to load in another simply because the codec is missing from that browser's media stack.
Extract a Frame Without Uploading
The full local workflow for turning a moment in a short video into a downloadable PNG is short enough to walk through in order. No server is involved, and no part of the file leaves the device during any of these steps.
- Choose one supported local video — an MP4, WebM, MOV, M4V, or Ogg file — and wait for its duration and preview to load inside the tool.
- Enter a frame time between zero and the displayed duration in the time field, using decimals when you need millisecond-level precision.
- Select Extract PNG frame, wait for the seek to complete, and inspect the still shown in the result panel.
- Download the local PNG; the filename reflects the requested time, and the pixel dimensions match the decoded frame exactly.
Two checks are worth doing before you treat the result as final. First, confirm that the time you entered actually points at the moment you wanted — the tool records the requested position, but the browser may land on a nearby decoded frame because of keyframes, variable frame rates, edit lists, timestamp rounding, and codec behavior. Second, confirm that the still you downloaded is a PNG of the expected dimensions; an unexpected size usually means the source file did not decode the way you assumed, and the tool will surface that as a failure rather than a blank image.
When Local Browser Extraction Is Not the Right Tool
A browser-based extractor handles a wide range of everyday tasks — thumbnails, slide references, presentation stills, quick references from media you own or may process — without ever uploading anything. There are still cases where it is the wrong fit, and the product contract calls them out directly so that expectations stay realistic.
Local extraction is not designed for frame-accurate editorial work, HDR or color-managed output, alpha workflows, batch extraction at exact frame numbers, long footage, or formats the browser cannot decode. Those scenarios need a dedicated desktop video tool with packet-level timestamp access, full codec support, and the ability to step through numbered frames rather than approximate media timeline positions. Within its bounded scope — a short, locally decodable clip, one still at a time, native PNG output — a browser extractor is the fastest way to confirm that nothing is uploaded while still producing a usable image.
For a deeper look, see Is My Video Uploaded When I Trim Video? A Local Answer.