[WAV] Read buffer: 380672 samples, 22050 Hz, 1 ch, 16 bit [Audio-Resample] 22050 Hz -> 16000 Hz, 380672 samples... [Audio-Resample] Done: 380672 -> 276225 samples ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97247 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97247 MiB load_backend: loaded CUDA backend from /mnt/workspace/git/qwenasr.cpp/build/libggml-cuda.so ggml_vulkan: Found 1 Vulkan devices: ggml_vulkan: 0 = NVIDIA RTX PRO 6000 Blackwell Workstation Edition (NVIDIA) | uma: 0 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2 load_backend: loaded Vulkan backend from /mnt/workspace/git/qwenasr.cpp/build/libggml-vulkan.so load_backend: loaded CPU backend from /mnt/workspace/git/qwenasr.cpp/build/libggml-cpu-zen4.so [Load] Pipeline backend: CUDA0 (CPU threads: 16) [GGUF] ../models/qwenasr-0.6B-Q4_K_M.gguf: 612 tensors, data at offset 5356096 [WeightCtx] Loaded 7 tensors, 19.5 MB into backend [ConvStem] Loaded: downsample_hidden 480, d_model 896 [WeightCtx] Loaded 294 tensors, 485.7 MB into backend [AudioEnc] Loaded: 18 layers, d_model 896, heads 14, head_dim 64, FFN 3584, output_dim 1024 [WeightCtx] Loaded 311 tensors, 456.1 MB into backend [Thinker] Loaded: 28 layers, hidden 1024, heads 16/8, head_dim 128, FFN 3072, RoPE theta 1000000, vocab 151936, tie_embd 1 [BPE] Loaded from GGUF: 151705 vocab, 151291 merges, eos_id=151643 [Dump] cpp/pcm16.bin [276225 floats] [Dump] cpp/mel.bin [220928 floats] [Dump] cpp/stem.bin [201600 floats] [Dump] cpp/windowed.bin [230400 floats] [KVCache] Allocated: 28 layers, 8 KV heads, head_dim 128, max_seq_len 755 -> 82 MB [Dump] cpp/hidden.bin [248832 floats] [Dump] cpp/logits.bin [151936 floats] ggml_backend_cuda_graph_compute: CUDA graph warmup complete ggml_backend_cuda_graph_compute: CUDA graph warmup reset ggml_backend_cuda_graph_compute: CUDA graph warmup complete [Perf] Resample 0.1 ms [Perf] Mel 4.6 ms [Perf] Tower 50.0 ms [Perf] Build 1.0 ms (embed + audio splice) [Perf] Prefill 10.4 ms [Perf] Decode 66.0 ms (52 tokens, 1.27 ms/tok, 788.3 tok/s) [Perf] Total 133.5 ms (audio 17.26 s, RTF 0.008) The following generation flags are not valid and may be ignored: ['temperature']. Set `TRANSFORMERS_VERBOSITY=info` for more details. Setting `pad_token_id` to `eos_token_id`:151645 for open-end generation. [GGML] Cmd: ../build/qwenasr-transcribe --model ../models/qwenasr-0.6B-Q4_K_M.gguf --file ../examples/freeman.wav --lang English --dump cpp -o cpp/asr-cpp.txt [GGML] Text: 'If you go into different cultures, they have different concepts of creation. They have their own creation story and of what an afterlife is, where you go, what you do, who you gonna be with. You know, people will say, "Well."' [Python] Text: 'If you go into different cultures, they have different concepts of creation. They have their own creation story and of what an afterlife is, where you go, what you do, who you going to be with. You know, people will say, "Well."' [Cossim] Mel cos 0.99999862 max_abs 3.780e-03 n 220928 [Cossim] Stem cos 0.99934707 max_abs 1.757e+00 n 201600 [Cossim] Windowed cos 0.99376007 max_abs 6.059e-02 n 230400 [Cossim] Hidden cos 0.93494992 max_abs 5.924e+01 n 248832 [Cossim] Text exact 0.00% norm_exact 0.00% word_ratio 0.963855