You write what you want to hear. The system plans the track, the art and the clip breakdown at the same table — with camera and photography learned from real clips already analyzed. Approve by phase, publish on the showcase, release on the channel.
Receives a phrase and returns music, cover and clip. Before spending anything, write a plano.json with the track structure, the art concept and the video breakdown — and only then execute, phase by phase, with your approval.
Watch a clip and return what the camera did: hex palette, lighting, lens, shot‑by‑shot with timecode, cuts per minute and prompts per scene. It is not a speech transcription — it is a reading of cinematography and editing direction.
It is the gateway: it receives the command in Telegram, creates the job, organizes the queue, triggers the worker and returns the result. It is through it that the two become a one‑line command, without terminal.
Planning is cheap, execution is expensive. All quality is decided before the first cent — that is why the product of musicavideo is not the file, it is the plano.json.
Asking a generic agent to “make a song about X” delivers a random track, a cover that does not match it, and no clip: the three come out disconnected because there was never a common plan. Here the structure of the song, the concept of the cover and the clip breakdown are decided together, looking at two measured material banks — musical styles (BPM, key, instrumentation, voice) and real visual references coming from analisevideo. Execution is just the consequence.
bash musicavideo.sh plano "turn‑around song, female rock,
about someone who builds in silence and now demands payment"
You send a phrase. Music, cover and clip come out — in phases, approving each part, and without spending anything until you tell it to spend. There is a local panel to listen, compare, send to trash and approve what goes up to the showcase.
bash analisevideo.sh analisa <url|arquivo>
bash analisevideo.sh prompts <slug>
Visual and cinematographic analysis with Gemini, archived in a searchable local database. It is not a speech transcription: what a cinematographer, an editor and a music producer see — hex palette, lighting scheme, apparent lens, shot‑by‑shot with timecode, cuts per minute, and 5 to 10 prompts per scene ready for Kling, Veo or Seedance.
/musicavideo turn‑around song, female rock,
about someone who builds in silence and now demands payment
You send the phrase in the chat. The bot creates the job, puts it in the queue, the worker executes and the files return to you in Telegram — without opening a terminal, without standing in front of the machine. The phases pause waiting for your approval and resume when you respond.
/analisevideo <link do clipe>
The clip is analyzed and enters the reference database — the same one the next track’s plan will consult. The two projects feed each other: what you analyze today improves what you produce tomorrow.
The local panel is where decisions are made. The showcase is only for reading — with one exception: the public like, which returns to the panel.
From modern sertanejo to Nordic folk and house — each with the original request, the lyrics, the prompts and the breakdown visible in the sheet.
Real clips analyzed frame‑by‑frame become the visual reference database that the plan consults.
The code is the URL of the page on the showcase — the production has a fixed address, it does not disappear in the feed.
The grid with the covers and the players, the sheet per MVD-000 with stacked versions and prompts, and the text analysis tab. The videos and tracks are served directly from the public dataset on Hugging Face.
It’s where productions are released in video, in short format. The showcase displays how each one was made; the channel shows it running.
The plan is the product. The music, the cover and the clip are what remain when the three decisions are made at the same table.
Showcase open, no login. Productions also appear on the channel @amoanimais2k.