Can ChatGPT analyze videos?
You cannot upload a video and have it watched. OpenAI's own FAQ answers it in one line: image inputs cannot handle videos and support static images only. The supported types are PNG, JPEG and non-animated GIF. What does work is different in kind — a live camera feed in Advanced Voice Mode, a YouTube page read as a web page, or frames you extract yourself and upload as images.
OpenAI gives this its own FAQ entry, and the answer is one sentence long.
Under the heading "Do the image inputs support videos?", the ChatGPT Image Inputs FAQ answers: "No, it cannot handle videos. It currently supports processing static images only." The supported file types are listed immediately after — PNG, JPEG, and non-animated GIF — which rules out the obvious workaround of converting a clip to an animated GIF.
The file uploads documentation agrees by omission. Its supported list covers text files, spreadsheets, presentations and documents — XLSX, XLS, CSV, TSV, DOCX, PPTX, PDF, TXT — with a 512 MB hard ceiling per file. No video container appears anywhere in it. So there are two separate upload paths into ChatGPT, and neither accepts video.
Live video is a genuinely different feature and it does exist: real-time camera and screen sharing in Advanced Voice Mode, subject to daily usage limits. Worth knowing that this is in flux — OpenAI's newer voice model, GPT-Live-1, does not support video or screen sharing at this time, and directs subscribers who need those to keep using Advanced Voice Mode. That is real-time perception, not analysis of a file you already have.
Scope: we did not run any of this — we hold no ChatGPT account, and everything here comes from OpenAI's current documentation with the dates it was last updated. The claim we are most confident about is the negative one, because OpenAI states it directly rather than leaving it to be inferred.
What to do instead.
If the information you need is visual and static — a slide, a diagram on a whiteboard, an error dialog, a chart that appears at 4:12 — take screenshots and upload those. ChatGPT accepts images on every plan including Free, and all ChatGPT models accept image inputs.
This fails for anything where the information is in the motion: judging a golf swing, diagnosing a stutter in a UI animation, following who moved where. Stills cannot carry that, and no amount of them will.
Be honest about what you actually want. "Summarise this video" almost always means "summarise what was said in it", and that is a text problem. Paste the transcript, or point ChatGPT at a page that carries one, and you get a better result than any frame analysis would give — faster, cheaper, and without the model guessing at pixels.
The same applies to YouTube links. With browsing available, ChatGPT can fetch the page and work from its title, description and any transcript published there. It is reading a web page, not watching a video, and the difference shows the moment the useful content is visual rather than spoken.
For "look at this and tell me what you see", Advanced Voice Mode's camera and screen share are the right tool — real-time, conversational, and subject to daily usage limits on all plans that have it.
Two caveats before you rely on it. It is not available in every region, and the newer GPT-Live-1 voice experience does not carry video or screen sharing yet, so which voice mode you land in determines whether the camera is there at all. And it cannot be pointed at a recording after the fact — the whole design assumes the thing is happening now.
Even on stills, OpenAI publishes an unusually candid list of where vision struggles, and it is worth reading before trusting a frame-by-frame analysis. The model is not suitable for specialised medical images such as CT scans; it does worse on non-Latin scripts; it may misread rotated or upside-down content; it struggles with graphs where meaning is carried by line style or colour; it is poor at precise spatial localisation; it gives approximate counts; and it struggles with panoramic and fisheye images.
Also relevant to screenshots specifically: images are resized before analysis, and file names and metadata are not processed at all. A timestamp burned into a frame may survive that resize; one in the filename will not be read.
Tools built to watch.
None of these is tested by us and we publish no quality comparison — they are named because they accept video where ChatGPT does not. Check current limits and terms with each vendor before building a workflow on one.
Frequently asked.
Quick follow-ups people search after this question.