Commit Graph
277 Commits
Author SHA1 Message Date
Leonid PershinandClaude Opus 5 7f441defcc feat: add start.sh for Linux/macOS, harden both launchers
start.sh mirrors start.bat: finds Python 3.10+, creates the venv,
installs PyTorch (CUDA index on Linux, the MPS build on macOS, or the CPU
index with --cpu), installs ACE-Step, reports the device and launches the
UI. Same flags, and the defaults can also come from the environment
(PORT=7870 ./start.sh). On macOS it passes --bf16 false, which the README
already calls for. Dropped start.sh from .gitignore, where it sat among
the upstream author's local scratch files.

Both launchers also gain two fixes found while testing on WSL:

- A venv is only accepted if pip works in it, not merely if the
  interpreter exists. A directory left by an interrupted install looked
  ready and then failed several steps later with a misleading "check your
  internet connection". Such a venv is now recreated, and if creation
  fails on Debian/Ubuntu the error points at python3-venv, which is the
  actual cause there.
- The launch banner announced the URL as if the server were already up,
  while model loading still had a minute to go. It now says the interface
  will be available once "Running on local URL" appears.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 20:31:18 +03:00
Leonid PershinandClaude Opus 5 e9ea6b9bab feat: generation options for infer.py, and split out training deps
infer.py could only render a random example from examples/input_params:
there was no way to pass your own prompt, lyrics, duration or seed, so
using it for anything specific meant writing a separate script. Add
--prompt, --lyrics/--lyrics_file, --duration, --steps, --guidance_scale,
--scheduler, --cfg_type, --omega_scale, --seed and --format alongside the
existing runtime flags. Without --prompt the old random-example behaviour
is kept, so existing invocations are unaffected.

Also move the training-only packages out of the default install.
datasets, pytorch_lightning, matplotlib, tensorboard and tensorboardX are
imported by trainer.py and convert2hf_dataset.py, never on the inference
path, but every user was installing them — and datasets==3.4.1 is a hard
pin that drags constraints onto huggingface-hub. They now live in
requirements-train.txt behind the existing (previously ineffective)
"train" extra: pip install -e ".[train]".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 20:31:07 +03:00
Leonid PershinandClaude Opus 5 c7953dc4e0 fix: don't crash the UI on output files with custom names
create_text2music_ui() sorted the saved *_input_params.json files with
int(name.split('_')[1]), which assumes the generated
output_<timestamp>_<idx>_ shape. Output names are user-controlled — via
infer.py --output_path or a save_path from the UI — so any other name
raised ValueError while the Blocks were being built and took the whole
interface down before it could start:

    ValueError: invalid literal for int() with base 10: 'base'

Sort by mtime instead. That is what "previous generated input params"
means anyway (newest first) and it works for any filename; a file that
disappears between listdir() and getmtime() sorts last rather than
raising.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 20:30:58 +03:00
Leonid PershinandClaude Opus 5 0584397884 fix: write and read audio with soundfile instead of torchaudio
torchaudio 2.11 routes torchaudio.save()/load() through TorchCodec and
ignores the `backend` argument, so every generation died at the save
step with "ImportError: TorchCodec is required for save_with_torchcodec"
after the diffusion had already finished. Reference-audio loading
(audio2audio, repaint, extend) and the training dataset loader hit the
same wall.

soundfile is already a required dependency and covers all four output
formats the UI offers, so use it directly rather than pulling in
TorchCodec and its native FFmpeg stack:

- pipeline_ace_step.save_wav_file(): sf.write(), transposing
  (channels, samples) -> (samples, channels); drops the now-unused
  torchaudio import
- MusicDCAE.load_audio() and text2music_dataset: sf.read(dtype=float32,
  always_2d=True), transposed back to (channels, samples)

torchaudio is still used for Resample/MelScale transforms, which are
unaffected.

Verified end to end: 10s generation on an RTX 3060 in 9.7s, output is
valid non-silent 48kHz stereo; load_audio round-trips it; wav/mp3/ogg/
flac all write and read back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 19:03:29 +03:00
Leonid PershinandClaude Opus 5 6e3273d049 deps: migrate the UI to Gradio 6
Gradio was unpinned, so a fresh install pulled Gradio 6, where
gr.Audio no longer accepts show_download_button — the UI crashed on
startup with TypeError. Move forward to Gradio 6 instead of pinning
back to 5:

- requirements: gradio>=6.0.0
- drop the removed show_download_button kwarg (4 call sites)
- requirements: transformers>=4.57.0 — Gradio 6 requires
  huggingface-hub>=1.0, which transformers==4.50.0 forbids. Only
  UMT5EncoderModel and AutoTokenizer are used, so the bump is safe.

Verified: pip check clean, UI renders and server callbacks work on
gradio 6.26.0 / transformers 5.16.1 / huggingface-hub 1.30.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 18:38:12 +03:00
Leonid PershinandClaude Opus 5 911899635a docs: add Russian README and a Windows start.bat launcher
- README_RU.md: full Russian translation, linked from README.md
- start.bat: one-click setup + launch for Windows. Finds Python 3.10+,
  creates the venv, installs PyTorch (CUDA or CPU) and ACE-Step, checks
  the GPU, then starts the web UI. Flags: --lowvram, --cpu, --share,
  --port, --device, --listen, --reinstall, --update, --setup, --help.
- Both READMEs document the launcher.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 18:38:05 +03:00
Gong Junmin 1bee4c9f5b Merge pull request #373 from iackov/fix/extend-mode-shape-mismatch 2026-02-15 12:57:54 +08:00
Gong Junmin 4a24cfb662 Merge pull request #375 from ace-step/copilot/add-external-link-for-model 2026-02-15 12:57:18 +08:00
copilot-swe-agent[bot]andChuxiJ 66b690b05b Add external link for ACE-Step v1.5 model
Co-authored-by: ChuxiJ <30956809+ChuxiJ@users.noreply.github.com>
2026-01-28 00:09:54 +00:00
copilot-swe-agent[bot] bf3510b949 Initial plan 2026-01-28 00:08:27 +00:00
jackj 435e9fd667 Fix: Shape mismatch in extend mode causing AssertionError
- Added automatic shape alignment for target_latents and x0
- Handles both shorter (padding) and longer (trimming) cases
- Fixes crash in extend mode with long audio files
- Minimal impact on audio quality (~0.05-0.15 sec)

Resolves issue where extend mode fails with AssertionError
when target_latents shape doesn't match x0 shape after
padding/trimming operations.
2026-01-23 23:15:42 +05:00
Gong Junmin 6ae0852b13 Merge pull request #282 from WebChatAppAi/main
Fix Incorrect Data Processing Flow in TRAIN_INSTRUCTION.md
2025-06-27 13:46:58 +08:00
WebChatAppAi 6e93273e81 Update TRAIN_INSTRUCTION.md
- Corrected file format requirements and naming conventions
2025-06-27 03:11:04 +06:00
WebChatAppAi faba128c5d Update TRAIN_INSTRUCTION.md 2025-06-27 03:09:42 +06:00
Gong Junmin d612523f50 Merge pull request #257 from SD-inst/ogg-fix 2025-06-04 08:46:37 +08:00
rkfg 9c5c92946d Use sox backend for ogg 2025-06-03 15:31:20 +03:00
Gong Junmin 0440ca24bb Merge pull request #256 from ace-step/update_tech_report_link
update tech report
2025-06-03 14:22:24 +08:00
chuxij 4a328731b8 update tech report 2025-06-03 06:21:26 +00:00
Gong Junmin 1028991adc Merge pull request #254 from ace-step/add_tech_report
add tech report
2025-06-02 11:05:06 +08:00
chuxij 40f010d9f7 add tech report 2025-06-02 03:03:32 +00:00
Sayo 4db9a5db12 Merge pull request #241 from ace-step/sayo_dev
[fix] export_quantized_weights import
2025-05-26 11:40:32 +08:00
Sayo 3e2cb3ee58 [fix] export_quantized_weights import 2025-05-26 03:40:01 +00:00
Gong Junmin f995fc2715 Merge pull request #233 from SamratBarai/patch-1 2025-05-24 09:58:57 +08:00
Samrat Barai 879e717c83 Update inference.ipynb to fix bug and improve UI
There was a problem in the script that caused two lines to be merged. Fixed the problem.

Also added some emojis to the cells
2025-05-23 17:52:35 +06:00
sean d2a13e52f4 Merge pull request #230 from ace-step/dev
fix bug caused by parameter positions in infer.py
2025-05-22 15:15:34 +08:00
sean 022a72881f fix bug caused by parameter postions in infer.py 2025-05-22 07:08:45 +00:00
Gong Junmin 6f5c0f88a7 Merge pull request #229 from SD-inst/lora-weight-fix 2025-05-22 09:46:38 +08:00
rkfg 05e641f6ec Fix lora reloading on weight change 2025-05-21 19:41:11 +03:00
Gong Junmin 2ba06ca1f8 Merge pull request #221 from SD-inst/lora-weight 2025-05-20 14:42:42 +08:00
rkfg fc984d81fa Add lora weight 2025-05-20 08:59:00 +03:00
Gong Junmin 969c6739bc Merge pull request #219 from craftmine1000/main 2025-05-20 07:26:54 +08:00
Michael Hedman 12bcab346e revert unlocking dependencies 2025-05-19 16:55:01 +02:00
Michael Hedman b4ea314e1c change normal notebook repo source to official repo 2025-05-19 15:49:00 +02:00
Michael Hedman 2c120c4808 change colab notebook repo source to official repo 2025-05-19 15:41:39 +02:00
Michael Hedman 9be6e4a905 Merge branch 'main' of https://github.com/ace-step/ACE-Step.git 2025-05-19 13:55:41 +02:00
Michael Hedman 1176d3bf97 unlock dependencies for future versions 2025-05-19 13:54:42 +02:00
Michael Hedman 730340f14a bump diffusers>=0.33.0 to prevent diffusers==0.32.2 dtype casting issue 2025-05-19 13:52:22 +02:00
Gong Junmin 516792aa82 Merge pull request #214 from craftmine1000/main 2025-05-19 17:50:15 +08:00
Gong Junmin f51063ee8b Merge pull request #215 from SD-inst/diffusers-fix 2025-05-19 17:48:52 +08:00
rkfg 481d950db4 Fix lora loading with diffusers 0.33.1 2025-05-19 09:33:56 +03:00
Michael Hedman bfac11fecf prepare for hub download of quantized models 2025-05-19 08:18:26 +02:00
Michael Hedman 401b910936 missed a dtype
cleanup immediate surroundings
2025-05-19 07:57:05 +02:00
Michael Hedman 0b379b6163 reduce device and dtype vars 2025-05-19 07:34:13 +02:00
Michael Hedman d470cd903f reduce path vars 2025-05-19 06:47:04 +02:00
Michael Hedman 21efcd3905 move language filters 2025-05-19 04:51:22 +02:00
Michael Hedman cdab63d688 remove fixed float16 COLAB_FREE env var
add ACE_PIPELINE_DTYPE env var dtype config
set colab notebook to float16
2025-05-19 04:17:38 +02:00
Michael Hedman e2efc727e9 update colab notebook 2025-05-19 03:31:40 +02:00
Michael Hedman 0e0021d754 Merge branch 'main' of https://github.com/ace-step/ACE-Step.git 2025-05-18 20:22:46 +02:00
Gong Junmin 74121a5d8a Merge pull request #209 from GeorgeDrb/memory-optimization 2025-05-18 21:24:42 +08:00
GeorgeDrb 375431486c Add memory cleanup to prevent VRAM overflow 2025-05-18 15:35:21 +03:00