Back to blog
PostBy MaikoAugust 2, 202613 min read

The Kumar Method: What It Is, Why It Went Viral, and How to Make One

A complete breakdown of the Kumar Method — the viral Instagram video format, where it came from, the five-line script, the editing rules that make it work, and how to make your own with AI.

kumar methodthe kumar methodkumar method generatorkumar method videokumar method scripthow to make a kumar method videokumar method aikumar method trendinstagram reelsviral video formatshort form videoai video generator

In June 2026 a man in a black turtleneck stood against a bright wall, let the light fall behind him until his face went half-dark, and calmly announced that he was a retired accountant who intended to steal everyone's job by becoming the biggest finance influencer in the world.

The video did tens of millions of views. Within weeks the structure had been lifted into every niche on the internet — dentists, dog groomers, junior developers, sourdough people. It picked up a name along the way: the Kumar Method.

This is a breakdown of what the format actually is, why it spread as fast as it did, the rules that separate the versions that land from the ones that don't, and how to make one yourself — by hand or with AI doing the expensive part.

What is the Kumar Method?

The Kumar Method is a short-form video format built on a single structural joke: cinematic treatment applied to an unremarkable claim.

Every version has the same parts:

  • A talking-head take. One person, filmed in portrait, delivering roughly five short lines straight down the lens with no expression at all.
  • Hard backlighting. A bright key behind or beside the subject so they read as a near-silhouette with a rim of light down one edge. Faces stay partly in shadow.
  • Cinematic cutaways. Studio shots of the same person — profile, arms crossed, seated leaning forward, hands in pockets, slow push-ins — cut between the spoken lines.
  • Giant red keyword type. One word from each line, punched onto the screen in enormous red letters, on the exact frame it is spoken.
  • A white flash. A single half-second ramp to white marking the turn from setup to threat.
  • A closing title card. The final line held for a couple of seconds in oversized type.

Nothing about that treatment matches what is being said, and that mismatch is the entire product. The Kumar Method takes the visual grammar of a film trailer and points it at someone announcing they do bookkeeping.

Where the Kumar Method came from

The format traces to the account @thekumarmethod, which posted its first video on Instagram and TikTok on 5 June 2026. The Instagram post accumulated tens of millions of views inside two weeks; the wider format has since driven well over a hundred million views across recreations.

The character is a retired accountant named Kumar with an implausible ambition and a completely straight face. He is presented as real, filmed as though the documentary crew showed up, and written as though the stakes are enormous.

Is Kumar a real person?

Short answer: treat "Kumar" as a character, not a biography.

The videos are performed, not documentary, and the account has been tied to a broader marketing campaign — including a website pitching the "engineered virality" system supposedly behind the character's rise. The Kumar accounts later distanced themselves from that site. Coverage since has treated the identity question as genuinely unresolved, and the account has not settled it.

That ambiguity is, arguably, part of why it travelled. A format you cannot quite tell is sincere is a format people argue about in the comments, and arguing in the comments is distribution.

Why the Kumar Method went viral

Plenty of formats look good and go nowhere. This one had six things going for it at once.

1. It is instantly legible on mute

Most people meet a Reel with the sound off. The red keyword cards mean the entire script is readable without audio — not as small burned-in subtitles, but as one word filling the screen. You can understand a Kumar Method video from the thumbnail rail.

2. The contrast does the comedic work, not the writing

You do not need a joke. You need a flat delivery and an ordinary job. The production supplies the tension; the script supplies the deflation. That means the format works for people who are not funny, which is most people, which is why so many were able to copy it successfully.

3. It is a fill-in-the-blank template

"This is a message to all ___. My name is ___. I'm a ___. And I'm gonna ___." The structure survives total substitution. Every niche on the internet could see their own version of it within about four seconds of watching the original.

4. It flatters the person making it

Most viral formats make you look silly. This one makes you look like the antagonist in a prestige drama. The cost of participating is unusually low because the downside — looking stupid on camera — has been engineered out. You look cool, and the joke is carried by the premise.

5. It is expensive-looking but cheap to fake

The cinematic shots read as a lighting setup, a location, and a crew. In practice they are a bright wall, a phone, and a hard light behind you — or, more often now, generated stills animated into four-second clips. The perceived production gap between an amateur version and the original is much smaller than in almost any other format.

6. The rhythm is fixed

The cuts land on a consistent grid rather than wherever the editor felt like. That regularity is what makes it feel professionally cut even when it isn't, and it's what makes it reproducible.

The Kumar Method script formula

Five lines. Always in this order.

BeatTemplateJob
1. The address"This is a message to all [niche]"Tell a specific group this is aimed at them
2. The name"My name is [your name]"Deliberately unremarkable — a beat of nothing
3. The ordinary truth"I'm a [ordinary job or status]"The contrast engine
4. The threat"And I'm gonna [take what's theirs] by becoming the biggest [thing] in the world"The absurd ambition, delivered like an ultimatum
5. The button"[Short closing question]"Becomes the title card

Some rules that matter more than they look:

Be specific in line one. "This is a message to all finance bros" works because finance bros exist as an identity people either claim or resent. "This is a message to everyone" works on nobody.

Line three should be mundane, not funny. Retired accountant. Substitute teacher. Part-time dog groomer. The moment you reach for something wacky — "I'm a professional balloon animal assassin" — you have written a sketch, and the cinematic treatment stops being ironic.

Keep line five to three words or fewer. It holds on screen as enormous type. "Do you get me now?" is four short ones and it barely fits. Long closers render as small text, which kills the card.

Never exceed five lines. The format's pacing is built around a very short runway. Six lines and the montage runs out of shots to cut against.

How to film a Kumar Method video

The shoot is the part people over-engineer. It should take fifteen minutes.

Framing. Phone in portrait, at eye level, roughly chest-up. Plain wall behind you. Nothing in shot that dates the video or explains you.

Lighting. This is the only technical decision that matters. Put your light behind or beside you, not in front. A window at your back, a lamp aimed at the wall, anything that leaves your face partly in shadow with a bright edge. A ring light pointed at your face produces a YouTube video, not a Kumar Method video.

Wardrobe. Dark, plain, no logos. The reference wears a black turtleneck for a reason: it disappears into the silhouette.

Delivery. Flat. Not angry, not amused, not doing a voice. Look down the lens. Let your face do almost nothing. If you can hear yourself performing, do it again quieter.

Pauses. Leave several seconds of silence between each line. This feels absurd while filming and is essential in the edit — it's what lets each line be cut apart and dropped onto its own beat without clipping words.

One take. Shoot the whole thing continuously rather than line by line. Your lighting, framing and energy stay identical, and the audio matches.

How to make the cinematic shots

This is where the work actually is, and where most manual attempts fall apart.

You need somewhere between five and ten cutaway shots of yourself: standing in silhouette, a rim-lit close-up, a side profile, seated leaning forward, arms crossed, hands in pockets, and ideally a fully-lit hero shot for the ending. In each one the camera should be moving very slowly — a push-in, an orbit, a lateral glide. Nothing else moves.

There are two routes.

Route one: shoot them

Same wall, same light. Move yourself and the phone through the poses, recording five to ten seconds of each. Free, and the identity is obviously perfect. The downside is that you need a slow, smooth camera move to get the cinematic feel, and handheld doesn't give you one.

Route two: generate them

Take one clear photo of yourself and use an image model to produce the poses, then an image-to-video model to animate each into a short clip with a slow camera move.

Two things reliably go wrong here:

  1. Face drift. Generate the shots one at a time and your face changes between them — slimmer here, different hairline there. Models "idealise" by default. Prompts have to explicitly forbid it, and even then, consistency across ten separate generations is the hard part.
  2. Too much motion. Asked to "animate," models make people walk, gesture, and emote. The format wants a slow blink and a breath. Prompts need to say subtle movement only and mean it.

One trick worth knowing: generating the shots as a single continuous multi-scene clip and then slicing it apart holds identity far better than ten separate generations, because the model is maintaining one subject across one output rather than re-inventing them ten times. It is also roughly a tenth of the cost.

How to edit it

The edit is where the format actually lives. Precise numbers, from frame-by-frame analysis of the reference clips:

  • Montage cuts land on a 0.9-second grid. Not "roughly a second" — a consistent interval. This regularity is most of the perceived polish.
  • One burst of three fast cuts at ~0.23 seconds, on the money shot only. Used once. Using it repeatedly turns the video into a generic hype edit.
  • The white flash is ~0.5 seconds, full frame, at the act break between setup and threat.
  • The closing card holds ~2.1 seconds. Long enough to screenshot.
  • Keywords are sized to fill about 90% of the frame width. This is why one-word beats read as enormous and long words read as small — it's fit-to-width, not a fixed font size.
  • Keyword colour is pure red (#FF0000), which is why it survives platform compression.

Keywords fire on the frame the word is spoken, not on the cut. Function words — "the", "a", "by" — mostly get dropped, though the reference is looser about this than you'd expect and does punch the occasional "BY".

Do this by hand and you are looking at an image tool, a video model, an editor, an auto-caption pass you then have to correct, and manual restyling of every keyword. It's an evening's work if the character stays consistent, and a weekend if it doesn't.

How AI automates the Kumar Method

Break the format into jobs and it's clear which half is worth automating:

JobWho should do it
Writing five linesYou. It's two minutes and it's the only creative decision.
Filming the takeYou. Your face and voice are the point.
Ten consistent cinematic shotsAI. This is the expensive, fiddly part.
Transcribing and timing the takeAI. Nobody should be nudging keyframes.
Picking keywords per lineAI, with your review.
Cutting to the grid, flash, card, musicAutomated. It's a fixed spec, not a taste decision.

The important distinction: AI should not be generating the performance. A Kumar Method video where the talking head is also synthetic loses the thing that makes the format land — a real person saying a ridiculous thing with a real face. Text-to-speech narration over generated footage is a different, much worse video.

Making one with Brainrot Shorts

Brainrot Shorts has a studio built specifically for this format rather than a general editor with a preset. The split follows the table above.

  1. Write your five lines. The studio shows the format's own template as placeholder text, in whichever of eleven languages you're filming in — English, Spanish, French, German, Portuguese, Italian, Dutch, Polish, Turkish, Indonesian, or Russian.
  2. Film one take. A built-in teleprompter paces each line to your language's reading speed, held deliberately slow, so there's nothing to memorise and no reason to rush.
  3. Upload one photo. The ten-shot cinematic bank is generated as a single continuous studio film and sliced apart — the consistency trick described above — so your face, outfit and lighting stay identical across every cutaway.
  4. Review and export. Check the line timing and keywords in a live preview before rendering, pick which shots make the cut, and export a clean 1080×1920 MP4 at 30fps with the music bed included and no watermark.

The cut grid, the burst, the white flash, the keyword sizing and the title card are all part of the renderer, because they're specification rather than preference. There is no text-to-speech anywhere in it — the audio is your take.

Should you still make one?

The format is now well past the point where the original script is usable. If you post "I'm a retired accountant" in your version, you have made a copy, and copies of a two-month-old format read as late.

What still works is the structure with genuinely local substitution: your actual niche, your actual unremarkable job, a threat that means something to the people you're addressing. The versions doing numbers now are the ones where the audience recognises the specific world being threatened.

The other honest use is commercial. The format was engineered to make an unremarkable person feel like a threat — which is, structurally, exactly what a small brand announcing itself needs to do. Swap the niche for your category and the ordinary job for whatever your founder actually is, and the joke works because it's true.

Ready to ship faster?

Turn what you learn here into clips, captions, and exports.

Brainrot Shorts is built for creators who want the posting volume of a media team without hiring one.

Keep reading

Related posts

PostJul 13, 202612 min read

AI Image & Video Generator: Every Model Explained

Generate AI images and videos with Nano Banana, GPT Image, FLUX, Sora, Veo, Kling, and more. Compare models and create better short-form content.

ai image generatorai video generatorimage to video aitext to video aigenerative mediaai brainrotnano bananasoraveoklingshort form video
Read post
PostJul 3, 20265 min read

How to Turn a Prompt Into a Video With AI (Free, Step by Step)

The complete prompt to video workflow: write one text prompt, pick a visual style and length, and let AI generate the script, scenes, cinematic motion, narration, and captions — plus the prompt-writing formula that separates good shorts from generic ones.

prompt to videoprompt to video aifree prompt to video aitext prompt to videoai video from text prompttext to video aiai video generatorfaceless shortsshort form video
Read post