Top 5 in AI

Guides

How to Turn Your City Into a GTA-Style AI Video (the 'GTA: Philly' Method, Reconstructed)

By the Top5Apps editorial team · Published October 2, 2026 · Updated October 2, 2026 · 7 min read

Share

Short answer: it's fake gameplay, not a mod — AI video prompted to look like a third-person open-world game set in a real city, cut together from 30-second generations. On October 2, YouTuber Jesse Wellens posted a two-minute clip captioned 'How to turn my hometown of Philadelphia into grand theft auto style video game,' with missions like 'Grab a cheesesteak at Pat's' and 'Meet Gilly at the UTV.' It's at about 270,000 views on X, and it never says how it was made; a reply asking 'what did you use to make it?' has gone unanswered. So this guide is a reconstruction, labeled as one. The method comes from five creators who did publish their prompts for the same format in September — all using Seedance 2.5 — and the recipe is consistent: lock the camera behind the player at shoulder height, let landmarks carry the city, give it one dumb local mission, and keep the HUD original. With GTA VI due November 19, this format has seven weeks of rising search ahead of it, and a rights holder that's already sending takedowns.

AI-generated fake gameplay frame of Philadelphia: a man in a grey hoodie seen from behind crossing the street toward Pat's King of Steaks at 9th and Wharton, with a mission box reading 'Grab a cheesesteak at Pat's,' a minimap in the lower left, and five wanted stars in the upper right
The whole formula in one frame: camera behind the player, a real corner, a mission card, a minimap, wanted stars. From Jesse Wellens' 'GTA: Philly' clip, shown editorially. Not affiliated with Rockstar Games.

What the clip is — and what we don't know about it

The post is a 2-minute-8-second video from Wellens, the former PrankvsPrank creator, and it's at least his second in the format — on September 14 he posted 'I turned my Burning Man experience into a video game similar to grand theft auto.' The look is consistent throughout: third-person camera, a mission box in the top-left corner, a minimap bottom-left, health and stamina bars, wanted stars. What's unconfirmed: the tools, the workflow, and the cameo viewers think they spotted. Search snippets of his TikTok suggest he offers followers 'the prompt to use in ChatGPT' for a comment, which is a lead magnet rather than a published method. Because a single Seedance generation tops out at 30 seconds, a 128-second video is necessarily several clips edited together — that much is arithmetic.

The documented version: five creators, one structure

ClipCreator, dateTool namedViews
Train-station diamond heist@saniaspeaks_, Sept 15Seedance 2.5 on OpenArt67,000
'GTA in Brazil' — deliver a package through a hillside neighborhood@john_my07, Sept 15Seedance 2.5 on Pollo AI43,000
Osaka grandma@SSSS_CRYPTOMAN, Sept 11Seedance 2.5, one 30-second generation37,000
Rural Japan chaos@AIwithkhan, Sept 14Seedance 2.5 on Pollo AI27,000
Old-man side mission (vertical)@RizwanAly07, Sept 14Seedance 2.5 on Sjinn22,000
Each post includes its full prompt. All are single 30-second generations with the HUD prompted inside the model, not added afterward. Several of these creators are promotional partners of the tools they name.

Read the five prompts side by side and the skeleton is identical: an opener asking for a '30-second ultra-realistic AAA third-person open-world gameplay sequence'; a main character block that points to an uploaded photo as the identity reference; timed beats (0–3 seconds, 3–8, 8–15, and so on) with the on-screen text for each quoted exactly; then blocks for visual style, camera, HUD, and negative requirements. The Brazil prompt's HUD block asks for 'a restrained GTA-inspired gameplay HUD integrated naturally into the screen: circular minimap, health bar, stamina bar, mission objective tracker, contextual interaction prompts.' Another adds the line that matters most for legibility: 'HUD must be minimal and event-based. English text only, exact spelling, no gibberish.'

A prompt template for your own city

Create a 30-second realistic third-person open-world gameplay sequence set in [CITY]. MAIN CHARACTER: the person in @Image1 — same face, hair, and outfit — seen from behind at shoulder height throughout. BEATS: 0–5s, the character walks toward [LANDMARK 1]; mission text top-left reads exactly "[MISSION LINE]". 5–15s, [one action: crossing the street, entering the shop, getting on the bus]. 15–25s, [arrival at LANDMARK 2, one obstacle]. 25–30s, mission text changes to "[COMPLETION LINE]". CITY DETAIL: [three real, specific things — a street corner, a transit vehicle, local signage, the team colors people wear]. CAMERA: fixed third-person follow camera, no drone shots, no slow motion, no shallow depth of field, slight handheld sway. HUD: original minimal game interface — square minimap bottom-left, two status bars beneath it, one mission box top-left. English text only, exact spelling. NEGATIVE: no real game logos or branding, no watermarks, no extra text, no cinematic cuts.
  • The camera is the trick. Behind the character, shoulder height, never cutting away. Drone shots, slow motion, and movie-style blur are what make AI video look like AI video; their absence is what makes it look like gameplay.
  • The city is carried by clutter, not a caption. Wellens' Philly is Pat's at 9th and Wharton, Eagles green in a stadium lot, a winged UTV. The Brazil clip is hillside houses, Portuguese signs, tangled wires. Name three real, specific things.
  • One dumb local mission. Deliver a package. Grab a cheesesteak. Meet someone at a vehicle. Not a plot.
  • One reference photo of the player, attached every time, so the character holds across clips — the same character-sheet discipline as our Jean Phil guide.
AI-generated fake gameplay frame of an Eagles tailgate: a player character seen from behind facing a man in a green Eagles hat, jacket, and shorts beside a green off-road vehicle with eagle wings painted on it, with a mission box reading 'Meet Gilly at the UTV,' a waypoint arrow, and a minimap
A mission beat: waypoint marker over the target, mission text in the corner, the lot's real details doing the rest. Note the team marks and sponsor signage — fine for a fan clip, a problem for anything commercial.

Getting the HUD text right: two routes

Route one, what every documented clip does: prompt the HUD in. Seedance 2.5 is unusually good at on-screen text — one detailed review scored its text rendering 4.5 out of 5, with lettering that 'stays legible and stable as the shot moves' — so quoting your mission line exactly and adding 'exact spelling' usually works. Route two, our recommendation when the words must be perfect or must change across a long edit: generate with clean corners and composite the interface afterward. Ask for 'no on-screen text, no HUD,' then add the mission box and minimap in CapCut or any editor. No published guide documents this for the format — it's our advice, on the same logic one Seedance guide gives for precise text: 'pair the model with prepared references and a bit of post-production.' It also solves the rights problem below, because you control exactly what the interface looks like.

Making it two minutes, and what that costs

  • Thirty seconds is the ceiling per generation, so a two-minute video is at least five clips — realistically ten to fifteen generations once you discard the bad takes.
  • Write it as missions. One mission per clip gives you natural cut points and hides the seams; a mission-complete card is a free transition.
  • Keep the player's outfit identical in every prompt and attach the same reference image, or the hoodie changes color between corners.
  • Budget: by Higgsfield's published rates a 15-second 1080p clip runs about $7 in credits; other Seedance hosts price differently and we couldn't confirm Pollo AI's per-clip cost. Draft at 480p or 720p and only render keepers at full resolution. Plan on tens of dollars for a first two-minute cut, not single digits.
  • Kling 3.0 and PixVerse can do the same look in shorter multi-shot clips (up to about 15 seconds); every published prompt we found for this format uses Seedance 2.5.

Rights: this is the one with an active enforcer

The minimap, the wanted stars, and the mission box are Rockstar's visual language, and Grand Theft Auto VI launches November 19, 2026, per Take-Two's August earnings release. Since July, according to games.gg, 'Rockstar Games has begun issuing DMCA takedowns against some of the more viral AI-generated GTA 6 posts on X.' The reported targets so far are fake GTA VI 'leaks' using Rockstar's characters and branding — not a parody of a real city — and we found no reaction to Wellens' clip. Rockstar's long-standing fan-content policy is generally understood to tolerate non-commercial fan videos of its games while reserving the right to remove anything; AI footage that merely looks like its games isn't clearly covered either way. Take-Two's CEO, for what it's worth, called the idea that AI tools let someone 'push a button and generate a hit' a 'laughable notion.'

  • Use an original HUD. One of the published prompts says it in so many words: 'Original realistic game HUD… No GTA branding/UI.' A square minimap and a plain mission box read as 'game' without copying a specific one.
  • Don't use the logo, the font, the title, or the characters. 'GTA-style' in your caption is description; a Grand Theft Auto wordmark in your video is a trademark.
  • Say 'fan-made, not affiliated with Rockstar Games.' Wellens' caption doesn't; yours should.
  • Label it AI. YouTube requires disclosure for content that 'alters footage of a real event or place' or 'generates a realistic scene that didn't actually occur' — and its exemption for 'gameplay footage from video games' doesn't cover footage that only imitates one.
  • Team logos, store signs, and real faces are other people's marks and likenesses. A fan clip of your city is one thing; an ad, a sponsored post, or a recognizable real person who didn't agree is another.

Our read

This is the most shareable of the formats we've covered, because the subject is the viewer's own hometown. Everyone wants to see their corner store with a mission marker over it, and the technique is just three disciplines — a locked camera, real local detail, exact on-screen text — applied to a model that happens to be good at all three. The gap between Wellens' caption and his content is the opportunity: he promised a how-to and shipped a showreel, and until someone publishes the stack the only honest version is the one above, assembled from creators who showed their work. The timing cuts both ways. Seven weeks before GTA VI is when search interest peaks, and when Rockstar's lawyers are most attentive. Make yours with an interface that's yours, say it's fan-made, and it's a love letter to your city. Copy theirs pixel for pixel and it's a takedown waiting for a view count.

Where these apps rank