working from /home/tony/Work/glockenspiel/sunoGPT working from /home/tony/Work/glockenspiel/sunoGPT working from /home/tony/Work/glockenspiel/sunoGPT [2024-05-16_04:33:37]: Failed to import xformers. [2024-05-16_04:33:37]: Failed to import flash_attn RMSNorm. Falling back to torch RMSNorm. Overriding: out_dir = /app/suno/checkpoints Overriding: data_dir = /app/suno/data/dpo/2b_before_recode_v0 Overriding: train_filename = data_tr.bin Overriding: train_metas_filename = meta_tr.jsonl Overriding: train_info_filename = info_tr.json Overriding: val_filename = data_val.bin Overriding: val_metas_filename = meta_val.jsonl Overriding: val_info_filename = info_val.json Overriding: learning_rate = 5e-07 Overriding: min_lr = 1e-08 Overriding: do_ipo = True Overriding: dpo_beta = 5.0 Overriding: semantic_codebook_weight = 4.0 Overriding: last_codebook_weight = 0.5 Overriding: warmup_iters = 200 Overriding: max_iters = 3000 Overriding: grad_clip = 0.1 Overriding: eval_interval = 500 Overriding: eval_iters = 25 Overriding: step_save_iters = 2000 Overriding: block_size = 8704 Overriding: t_text = 2560 Overriding: t_memmap = 6016 Overriding: t_audio = 6144 Overriding: use_rotary_pos_emb = True Overriding: rope_theta = 500000 Overriding: use_qk_norm = False Overriding: activation_f = gelu Overriding: n_layer = 40 Overriding: n_head = 40 Overriding: d_head = 128 Overriding: n_kv_head = 8 Overriding: attention_type = tao Overriding: gradient_accumulation_steps = 1 Overriding: batch_size = 2 Overriding: fsdp = True Overriding: sharding_strategy = full_shard Overriding: grad_checkpointing = True Overriding: preload_checkpoint = /app/suno/data/dpo/models/model_13b_full.pt Overriding: preload_strict = False Overriding: local_cache_dir = /mnt/localdisk/tmp_tony Overriding: wandb_log = True Overriding: wandb_project = chirp-v4-dpo Overriding: wandb_run_name = dpo_13b_v0 [2024-05-16_04:33:39]: ddp init, rank 16, local_rank 0 [2024-05-16_04:33:39]: ddp init, rank 17, local_rank 1 [2024-05-16_04:33:39]: ddp init, rank 18, local_rank 2 [2024-05-16_04:33:39]: ddp init, rank 20, local_rank 4 [2024-05-16_04:33:39]: ddp init, rank 0, local_rank 0 [2024-05-16_04:33:39]: ddp init, rank 21, local_rank 5 [2024-05-16_04:33:39]: ddp init, rank 5, local_rank 5 [2024-05-16_04:33:39]: ddp init, rank 6, local_rank 6 [2024-05-16_04:33:39]: ddp init, rank 1, local_rank 1 [2024-05-16_04:33:39]: ddp init, rank 3, local_rank 3 [2024-05-16_04:33:39]: ddp init, rank 7, local_rank 7 [2024-05-16_04:33:39]: ddp init, rank 2, local_rank 2 [2024-05-16_04:33:39]: ddp init, rank 4, local_rank 4 NCCL version 2.20.5+cuda12.4 [2024-05-16_04:33:39]: ddp init, rank 19, local_rank 3 [2024-05-16_04:33:39]: ddp init, rank 23, local_rank 7 [2024-05-16_04:33:39]: ddp init, rank 22, local_rank 6 [2024-05-16_04:33:40]: ddp init, rank 8, local_rank 0 [2024-05-16_04:33:40]: ddp init, rank 9, local_rank 1 [2024-05-16_04:33:40]: ddp init, rank 13, local_rank 5 [2024-05-16_04:33:40]: ddp init, rank 10, local_rank 2 [2024-05-16_04:33:40]: ddp init, rank 15, local_rank 7 [2024-05-16_04:33:40]: ddp init, rank 14, local_rank 6 [2024-05-16_04:33:40]: ddp init, rank 11, local_rank 3 [2024-05-16_04:33:40]: ddp init, rank 12, local_rank 4 [2024-05-16_04:33:52]: loss discounts for codebooks: [0.308 0.077 0.073 0.07 0.066 0.063 0.059 0.056 0.052 0.049 0.045 0.042 0.038] [2024-05-16_04:33:58]: logging checkpoint here: /app/suno/checkpoints/2024-05-16_04-33-58 [2024-05-16_04:33:58]: loading data... [2024-05-16_04:33:58]: indexed 100.0% of data [2024-05-16_04:33:58]: 1,568 lines of data_val.bin loaded. [2024-05-16_04:34:00]: indexed 100.0% of data [2024-05-16_04:34:00]: 155,220 lines of data_tr.bin loaded. [2024-05-16_04:34:00]: train data weights: 50.0% perference_0 50.0% perference_1 [2024-05-16_04:34:00]: done loading data [2024-05-16_04:34:00]: Initializing a new model from scratch [2024-05-16_04:35:35]: number of parameters: 11363M [2024-05-16_04:37:10]: number of parameters: 11363M [2024-05-16_04:37:10]: not compiling model. [2024-05-16_04:37:10]: verifying model args... [2024-05-16_04:37:10]: careful, using approximation for checkpoint loading. could be wrong in principle [2024-05-16_04:37:25]: careful, using approximation for checkpoint loading. could be wrong in principle [2024-05-16_04:37:25]: loading model state_dict on gpu 0 [2024-05-16_04:37:40]: loading model state_dict on gpu 1 [2024-05-16_04:37:56]: loading model state_dict on gpu 2 [2024-05-16_04:38:12]: loading model state_dict on gpu 3 [2024-05-16_04:38:28]: loading model state_dict on gpu 4 [2024-05-16_04:38:44]: loading model state_dict on gpu 5 [2024-05-16_04:39:00]: loading model state_dict on gpu 6 [2024-05-16_04:39:15]: loading model state_dict on gpu 7 [2024-05-16_04:39:31]: verifying model args... [2024-05-16_04:39:31]: careful, using approximation for checkpoint loading. could be wrong in principle [2024-05-16_04:39:43]: careful, using approximation for checkpoint loading. could be wrong in principle [2024-05-16_04:39:43]: loading model state_dict on gpu 0 [2024-05-16_04:39:59]: loading model state_dict on gpu 1 [2024-05-16_04:40:15]: loading model state_dict on gpu 2 [2024-05-16_04:40:30]: loading model state_dict on gpu 3 [2024-05-16_04:40:46]: loading model state_dict on gpu 4 [2024-05-16_04:41:02]: loading model state_dict on gpu 5 [2024-05-16_04:41:17]: loading model state_dict on gpu 6 [2024-05-16_04:41:33]: loading model state_dict on gpu 7 [2024-05-16_04:41:48]: wrapping model in FSDP .... [2024-05-16_04:42:11]: applying fsdp activation checkpointing... [2024-05-16_04:42:11]: num decayed parameter tensors: 267, with 479,536,640 parameters [2024-05-16_04:42:11]: num non-decayed parameter tensors: 84, with 204,800 parameters [2024-05-16_04:42:11]: using fused Optimizer: False [2024-05-16_04:42:11]: model setup done [2024-05-16_04:42:12]: training... [2024-05-16_04:43:26]: loss estimation took 74.4 seconds. (100.0% of loop) [2024-05-16_04:43:26]: step 0: train loss 4.3086, val loss 4.3707 [2024-05-16_04:43:33]: iter 0: avg_loss 0.010, avg_acc 0.000, step_time 80861.4ms, mfu 0.0%, throughput 0k tok/s, total time 81s [2024-05-16_04:45:32]: iter 25: avg_loss 0.010, avg_acc 0.000, step_time 4712.3ms, mfu 106.1%, throughput 89k tok/s, total time 200s [2024-05-16_04:47:31]: iter 50: avg_loss 0.010, avg_acc 0.000, step_time 4714.2ms, mfu 106.0%, throughput 89k tok/s, total time 320s [2024-05-16_04:49:31]: iter 75: avg_loss 0.010, avg_acc 0.000, step_time 4710.2ms, mfu 106.1%, throughput 89k tok/s, total time 439s [2024-05-16_04:51:31]: iter 100: avg_loss 0.011, avg_acc 0.080, step_time 4714.6ms, mfu 106.0%, throughput 89k tok/s, total time 559s [2024-05-16_04:53:30]: iter 125: avg_loss 0.010, avg_acc 0.040, step_time 4721.5ms, mfu 105.8%, throughput 88k tok/s, total time 678s [2024-05-16_04:55:30]: iter 150: avg_loss 0.010, avg_acc 0.280, step_time 4711.1ms, mfu 106.1%, throughput 89k tok/s, total time 798s [2024-05-16_04:57:29]: iter 175: avg_loss 0.011, avg_acc 0.160, step_time 4711.8ms, mfu 106.1%, throughput 89k tok/s, total time 918s [2024-05-16_04:59:29]: iter 200: avg_loss 0.010, avg_acc 0.360, step_time 4713.0ms, mfu 106.0%, throughput 89k tok/s, total time 1037s [2024-05-16_05:01:29]: iter 225: avg_loss 0.011, avg_acc 0.280, step_time 4714.3ms, mfu 106.0%, throughput 89k tok/s, total time 1157s [2024-05-16_05:03:29]: iter 250: avg_loss 0.008, avg_acc 0.560, step_time 4723.0ms, mfu 105.8%, throughput 88k tok/s, total time 1277s [2024-05-16_05:05:28]: iter 275: avg_loss 0.008, avg_acc 0.480, step_time 4730.1ms, mfu 105.7%, throughput 88k tok/s, total time 1397s [2024-05-16_05:07:29]: iter 300: avg_loss 0.009, avg_acc 0.520, step_time 4725.1ms, mfu 105.8%, throughput 88k tok/s, total time 1517s [2024-05-16_05:09:28]: iter 325: avg_loss 0.009, avg_acc 0.400, step_time 4726.7ms, mfu 105.7%, throughput 88k tok/s, total time 1636s [2024-05-16_05:11:28]: iter 350: avg_loss 0.010, avg_acc 0.440, step_time 4718.5ms, mfu 105.9%, throughput 89k tok/s, total time 1756s [2024-05-16_05:13:27]: iter 375: avg_loss 0.011, avg_acc 0.400, step_time 4731.6ms, mfu 105.6%, throughput 88k tok/s, total time 1876s [2024-05-16_05:15:27]: iter 400: avg_loss 0.007, avg_acc 0.640, step_time 4713.6ms, mfu 106.0%, throughput 89k tok/s, total time 1995s [2024-05-16_05:17:27]: iter 425: avg_loss 0.011, avg_acc 0.400, step_time 4719.8ms, mfu 105.9%, throughput 89k tok/s, total time 2115s [2024-05-16_05:19:27]: iter 450: avg_loss 0.010, avg_acc 0.440, step_time 4723.8ms, mfu 105.8%, throughput 88k tok/s, total time 2235s [2024-05-16_05:21:26]: iter 475: avg_loss 0.009, avg_acc 0.560, step_time 4719.8ms, mfu 105.9%, throughput 89k tok/s, total time 2354s [2024-05-16_05:24:30]: loss estimation took 68.8 seconds. (2.8% of loop) [2024-05-16_05:24:30]: step 500: train loss 4.2893, val loss 4.3052 [2024-05-16_05:25:59]: saving checkpoint to /app/suno/checkpoints/2024-05-16_04-33-58 [2024-05-16_05:34:59]: saving took 629.0 seconds. (25.5% of loop) [2024-05-16_05:35:10]: iter 500: avg_loss 0.010, avg_acc 0.520, step_time 708479.8ms, mfu 0.7%, throughput 1k tok/s, total time 3178s [2024-05-16_05:37:09]: iter 525: avg_loss 0.011, avg_acc 0.440, step_time 4728.9ms, mfu 105.7%, throughput 88k tok/s, total time 3297s [2024-05-16_05:39:09]: iter 550: avg_loss 0.009, avg_acc 0.520, step_time 4720.5ms, mfu 105.9%, throughput 89k tok/s, total time 3417s [2024-05-16_05:41:09]: iter 575: avg_loss 0.010, avg_acc 0.520, step_time 4713.2ms, mfu 106.0%, throughput 89k tok/s, total time 3537s [2024-05-16_05:43:09]: iter 600: avg_loss 0.010, avg_acc 0.480, step_time 5079.6ms, mfu 98.4%, throughput 82k tok/s, total time 3657s [2024-05-16_05:45:09]: iter 625: avg_loss 0.008, avg_acc 0.640, step_time 4717.9ms, mfu 105.9%, throughput 89k tok/s, total time 3777s [2024-05-16_05:47:09]: iter 650: avg_loss 0.008, avg_acc 0.560, step_time 4714.6ms, mfu 106.0%, throughput 89k tok/s, total time 3897s [2024-05-16_05:49:09]: iter 675: avg_loss 0.011, avg_acc 0.400, step_time 4726.3ms, mfu 105.7%, throughput 88k tok/s, total time 4017s [2024-05-16_05:51:08]: iter 700: avg_loss 0.010, avg_acc 0.520, step_time 4730.3ms, mfu 105.7%, throughput 88k tok/s, total time 4136s [2024-05-16_05:53:08]: iter 725: avg_loss 0.010, avg_acc 0.480, step_time 4714.6ms, mfu 106.0%, throughput 89k tok/s, total time 4256s [2024-05-16_05:55:08]: iter 750: avg_loss 0.008, avg_acc 0.640, step_time 4722.3ms, mfu 105.8%, throughput 88k tok/s, total time 4376s [2024-05-16_05:57:08]: iter 775: avg_loss 0.009, avg_acc 0.440, step_time 4734.4ms, mfu 105.6%, throughput 88k tok/s, total time 4496s [2024-05-16_05:59:07]: iter 800: avg_loss 0.008, avg_acc 0.480, step_time 4714.5ms, mfu 106.0%, throughput 89k tok/s, total time 4615s [2024-05-16_06:01:07]: iter 825: avg_loss 0.010, avg_acc 0.440, step_time 4715.6ms, mfu 106.0%, throughput 89k tok/s, total time 4735s [2024-05-16_06:03:07]: iter 850: avg_loss 0.010, avg_acc 0.440, step_time 4717.8ms, mfu 105.9%, throughput 89k tok/s, total time 4855s [2024-05-16_06:05:07]: iter 875: avg_loss 0.010, avg_acc 0.560, step_time 4731.2ms, mfu 105.6%, throughput 88k tok/s, total time 4975s [2024-05-16_06:07:06]: iter 900: avg_loss 0.008, avg_acc 0.640, step_time 4727.7ms, mfu 105.7%, throughput 88k tok/s, total time 5095s [2024-05-16_06:09:06]: iter 925: avg_loss 0.008, avg_acc 0.560, step_time 4727.3ms, mfu 105.7%, throughput 88k tok/s, total time 5214s [2024-05-16_06:11:06]: iter 950: avg_loss 0.011, avg_acc 0.400, step_time 4930.8ms, mfu 101.4%, throughput 85k tok/s, total time 5334s [2024-05-16_06:13:06]: iter 975: avg_loss 0.009, avg_acc 0.600, step_time 4742.5ms, mfu 105.4%, throughput 88k tok/s, total time 5454s [2024-05-16_06:16:11]: loss estimation took 69.4 seconds. (2.2% of loop) [2024-05-16_06:16:11]: step 1000: train loss 4.3687, val loss 4.2931 [2024-05-16_06:17:39]: saving checkpoint to /app/suno/checkpoints/2024-05-16_04-33-58 [2024-05-16_06:30:34]: saving took 863.1 seconds. (27.8% of loop) [2024-05-16_06:30:45]: iter 1000: avg_loss 0.008, avg_acc 0.600, step_time 943337.7ms, mfu 0.5%, throughput 0k tok/s, total time 6513s [2024-05-16_06:32:45]: iter 1025: avg_loss 0.008, avg_acc 0.520, step_time 4750.9ms, mfu 105.2%, throughput 88k tok/s, total time 6633s [2024-05-16_06:34:46]: iter 1050: avg_loss 0.010, avg_acc 0.480, step_time 4747.3ms, mfu 105.3%, throughput 88k tok/s, total time 6754s [2024-05-16_06:36:46]: iter 1075: avg_loss 0.006, avg_acc 0.760, step_time 4722.6ms, mfu 105.8%, throughput 88k tok/s, total time 6874s [2024-05-16_06:38:47]: iter 1100: avg_loss 0.008, avg_acc 0.680, step_time 4734.4ms, mfu 105.6%, throughput 88k tok/s, total time 6995s [2024-05-16_06:40:47]: iter 1125: avg_loss 0.010, avg_acc 0.520, step_time 4760.1ms, mfu 105.0%, throughput 88k tok/s, total time 7115s [2024-05-16_06:42:47]: iter 1150: avg_loss 0.009, avg_acc 0.560, step_time 4752.1ms, mfu 105.2%, throughput 88k tok/s, total time 7235s [2024-05-16_06:44:47]: iter 1175: avg_loss 0.007, avg_acc 0.720, step_time 4747.7ms, mfu 105.3%, throughput 88k tok/s, total time 7355s [2024-05-16_06:46:47]: iter 1200: avg_loss 0.009, avg_acc 0.600, step_time 4739.7ms, mfu 105.4%, throughput 88k tok/s, total time 7475s [2024-05-16_06:48:47]: iter 1225: avg_loss 0.008, avg_acc 0.600, step_time 4944.1ms, mfu 101.1%, throughput 85k tok/s, total time 7595s [2024-05-16_06:50:47]: iter 1250: avg_loss 0.008, avg_acc 0.720, step_time 4765.0ms, mfu 104.9%, throughput 88k tok/s, total time 7715s [2024-05-16_06:52:46]: iter 1275: avg_loss 0.007, avg_acc 0.680, step_time 4752.6ms, mfu 105.2%, throughput 88k tok/s, total time 7835s [2024-05-16_06:54:47]: iter 1300: avg_loss 0.009, avg_acc 0.480, step_time 4763.1ms, mfu 104.9%, throughput 88k tok/s, total time 7955s [2024-05-16_06:56:47]: iter 1325: avg_loss 0.007, avg_acc 0.720, step_time 4753.8ms, mfu 105.1%, throughput 88k tok/s, total time 8075s [2024-05-16_06:58:47]: iter 1350: avg_loss 0.007, avg_acc 0.720, step_time 4761.0ms, mfu 105.0%, throughput 88k tok/s, total time 8195s [2024-05-16_07:00:47]: iter 1375: avg_loss 0.009, avg_acc 0.600, step_time 4745.7ms, mfu 105.3%, throughput 88k tok/s, total time 8315s [2024-05-16_07:02:47]: iter 1400: avg_loss 0.006, avg_acc 0.800, step_time 4757.9ms, mfu 105.0%, throughput 88k tok/s, total time 8436s [2024-05-16_07:04:48]: iter 1425: avg_loss 0.010, avg_acc 0.480, step_time 4959.3ms, mfu 100.8%, throughput 84k tok/s, total time 8556s [2024-05-16_07:06:48]: iter 1450: avg_loss 0.008, avg_acc 0.520, step_time 4730.3ms, mfu 105.7%, throughput 88k tok/s, total time 8676s [2024-05-16_07:08:47]: iter 1475: avg_loss 0.007, avg_acc 0.680, step_time 4751.6ms, mfu 105.2%, throughput 88k tok/s, total time 8796s [2024-05-16_07:11:52]: loss estimation took 69.2 seconds. (2.1% of loop) [2024-05-16_07:11:52]: step 1500: train loss 4.3491, val loss 4.3162 [2024-05-16_07:13:20]: saving checkpoint to /app/suno/checkpoints/2024-05-16_04-33-58 [2024-05-16_07:20:18]: saving took 505.8 seconds. (15.1% of loop) [2024-05-16_07:20:29]: iter 1500: avg_loss 0.009, avg_acc 0.560, step_time 585701.5ms, mfu 0.9%, throughput 1k tok/s, total time 9497s [2024-05-16_07:22:29]: iter 1525: avg_loss 0.010, avg_acc 0.640, step_time 5056.2ms, mfu 98.8%, throughput 83k tok/s, total time 9617s [2024-05-16_07:24:29]: iter 1550: avg_loss 0.009, avg_acc 0.560, step_time 4938.1ms, mfu 101.2%, throughput 85k tok/s, total time 9737s [2024-05-16_07:26:28]: iter 1575: avg_loss 0.007, avg_acc 0.680, step_time 4725.5ms, mfu 105.8%, throughput 88k tok/s, total time 9856s [2024-05-16_07:28:28]: iter 1600: avg_loss 0.009, avg_acc 0.600, step_time 4725.7ms, mfu 105.8%, throughput 88k tok/s, total time 9976s [2024-05-16_07:30:28]: iter 1625: avg_loss 0.005, avg_acc 0.800, step_time 4944.5ms, mfu 101.1%, throughput 84k tok/s, total time 10096s [2024-05-16_07:32:27]: iter 1650: avg_loss 0.006, avg_acc 0.720, step_time 4732.5ms, mfu 105.6%, throughput 88k tok/s, total time 10215s [2024-05-16_07:34:26]: iter 1675: avg_loss 0.005, avg_acc 0.720, step_time 4733.1ms, mfu 105.6%, throughput 88k tok/s, total time 10335s [2024-05-16_07:36:26]: iter 1700: avg_loss 0.007, avg_acc 0.640, step_time 4938.6ms, mfu 101.2%, throughput 85k tok/s, total time 10454s [2024-05-16_07:38:26]: iter 1725: avg_loss 0.009, avg_acc 0.520, step_time 4746.4ms, mfu 105.3%, throughput 88k tok/s, total time 10574s [2024-05-16_07:40:25]: iter 1750: avg_loss 0.007, avg_acc 0.720, step_time 4754.0ms, mfu 105.1%, throughput 88k tok/s, total time 10693s [2024-05-16_07:42:26]: iter 1775: avg_loss 0.010, avg_acc 0.520, step_time 4730.4ms, mfu 105.6%, throughput 88k tok/s, total time 10814s [2024-05-16_07:44:25]: iter 1800: avg_loss 0.007, avg_acc 0.600, step_time 4728.8ms, mfu 105.7%, throughput 88k tok/s, total time 10933s [2024-05-16_07:46:28]: iter 1825: avg_loss 0.008, avg_acc 0.640, step_time 7717.5ms, mfu 64.8%, throughput 54k tok/s, total time 11056s [2024-05-16_07:48:28]: iter 1850: avg_loss 0.009, avg_acc 0.480, step_time 4736.8ms, mfu 105.5%, throughput 88k tok/s, total time 11176s [2024-05-16_07:50:28]: iter 1875: avg_loss 0.010, avg_acc 0.560, step_time 4741.3ms, mfu 105.4%, throughput 88k tok/s, total time 11296s [2024-05-16_07:52:27]: iter 1900: avg_loss 0.012, avg_acc 0.480, step_time 4734.2ms, mfu 105.6%, throughput 88k tok/s, total time 11415s [2024-05-16_07:54:27]: iter 1925: avg_loss 0.008, avg_acc 0.640, step_time 4735.5ms, mfu 105.5%, throughput 88k tok/s, total time 11535s [2024-05-16_07:56:27]: iter 1950: avg_loss 0.012, avg_acc 0.280, step_time 4728.7ms, mfu 105.7%, throughput 88k tok/s, total time 11655s [2024-05-16_07:58:27]: iter 1975: avg_loss 0.006, avg_acc 0.760, step_time 4729.8ms, mfu 105.7%, throughput 88k tok/s, total time 11775s [2024-05-16_08:01:34]: loss estimation took 68.9 seconds. (2.3% of loop) [2024-05-16_08:01:34]: step 2000: train loss 4.4054, val loss 4.2996 [2024-05-16_08:03:00]: saving checkpoint to /app/suno/checkpoints/2024-05-16_04-33-58 [2024-05-16_08:14:54]: saving took 800.9 seconds. (26.9% of loop) [2024-05-16_08:15:06]: iter 2000: avg_loss 0.008, avg_acc 0.520, step_time 880968.2ms, mfu 0.6%, throughput 0k tok/s, total time 12774s [2024-05-16_08:17:05]: iter 2025: avg_loss 0.007, avg_acc 0.560, step_time 4746.6ms, mfu 105.3%, throughput 88k tok/s, total time 12893s [2024-05-16_08:19:05]: iter 2050: avg_loss 0.007, avg_acc 0.720, step_time 5086.0ms, mfu 98.3%, throughput 82k tok/s, total time 13013s [2024-05-16_08:21:04]: iter 2075: avg_loss 0.006, avg_acc 0.760, step_time 4758.7ms, mfu 105.0%, throughput 88k tok/s, total time 13133s [2024-05-16_08:23:04]: iter 2100: avg_loss 0.009, avg_acc 0.640, step_time 4725.2ms, mfu 105.8%, throughput 88k tok/s, total time 13252s [2024-05-16_08:25:04]: iter 2125: avg_loss 0.007, avg_acc 0.640, step_time 4738.4ms, mfu 105.5%, throughput 88k tok/s, total time 13372s [2024-05-16_08:27:04]: iter 2150: avg_loss 0.007, avg_acc 0.560, step_time 4743.3ms, mfu 105.4%, throughput 88k tok/s, total time 13492s [2024-05-16_08:29:03]: iter 2175: avg_loss 0.008, avg_acc 0.560, step_time 4752.3ms, mfu 105.2%, throughput 88k tok/s, total time 13612s [2024-05-16_08:31:03]: iter 2200: avg_loss 0.006, avg_acc 0.720, step_time 4739.4ms, mfu 105.4%, throughput 88k tok/s, total time 13731s [2024-05-16_08:33:03]: iter 2225: avg_loss 0.007, avg_acc 0.600, step_time 4739.1ms, mfu 105.5%, throughput 88k tok/s, total time 13851s [2024-05-16_08:35:03]: iter 2250: avg_loss 0.007, avg_acc 0.680, step_time 4747.2ms, mfu 105.3%, throughput 88k tok/s, total time 13971s [2024-05-16_08:37:02]: iter 2275: avg_loss 0.009, avg_acc 0.640, step_time 4737.0ms, mfu 105.5%, throughput 88k tok/s, total time 14090s [2024-05-16_08:39:02]: iter 2300: avg_loss 0.009, avg_acc 0.560, step_time 4719.5ms, mfu 105.9%, throughput 89k tok/s, total time 14210s [2024-05-16_08:41:01]: iter 2325: avg_loss 0.007, avg_acc 0.760, step_time 4719.6ms, mfu 105.9%, throughput 89k tok/s, total time 14329s [2024-05-16_08:43:01]: iter 2350: avg_loss 0.009, avg_acc 0.560, step_time 4720.1ms, mfu 105.9%, throughput 89k tok/s, total time 14449s [2024-05-16_08:45:00]: iter 2375: avg_loss 0.007, avg_acc 0.760, step_time 4928.9ms, mfu 101.4%, throughput 85k tok/s, total time 14568s [2024-05-16_08:47:00]: iter 2400: avg_loss 0.005, avg_acc 0.760, step_time 4708.3ms, mfu 106.1%, throughput 89k tok/s, total time 14688s [2024-05-16_08:48:59]: iter 2425: avg_loss 0.007, avg_acc 0.720, step_time 4721.4ms, mfu 105.8%, throughput 88k tok/s, total time 14807s [2024-05-16_08:50:59]: iter 2450: avg_loss 0.006, avg_acc 0.840, step_time 4939.6ms, mfu 101.2%, throughput 85k tok/s, total time 14927s [2024-05-16_08:52:58]: iter 2475: avg_loss 0.006, avg_acc 0.720, step_time 4749.6ms, mfu 105.2%, throughput 88k tok/s, total time 15046s [2024-05-16_08:56:02]: loss estimation took 69.1 seconds. (2.1% of loop) [2024-05-16_08:56:02]: step 2500: train loss 4.3480, val loss 4.3362 [2024-05-16_08:57:30]: saving checkpoint to /app/suno/checkpoints/2024-05-16_04-33-58 [2024-05-16_09:04:38]: saving took 515.6 seconds. (15.8% of loop) [2024-05-16_09:04:49]: iter 2500: avg_loss 0.007, avg_acc 0.720, step_time 595902.4ms, mfu 0.8%, throughput 1k tok/s, total time 15757s [2024-05-16_09:06:48]: iter 2525: avg_loss 0.006, avg_acc 0.800, step_time 4740.0ms, mfu 105.4%, throughput 88k tok/s, total time 15876s [2024-05-16_09:08:48]: iter 2550: avg_loss 0.007, avg_acc 0.640, step_time 4747.2ms, mfu 105.3%, throughput 88k tok/s, total time 15996s [2024-05-16_09:10:47]: iter 2575: avg_loss 0.009, avg_acc 0.440, step_time 4928.2ms, mfu 101.4%, throughput 85k tok/s, total time 16115s [2024-05-16_09:12:47]: iter 2600: avg_loss 0.007, avg_acc 0.760, step_time 4760.4ms, mfu 105.0%, throughput 88k tok/s, total time 16235s [2024-05-16_09:14:47]: iter 2625: avg_loss 0.008, avg_acc 0.760, step_time 4742.9ms, mfu 105.4%, throughput 88k tok/s, total time 16355s [2024-05-16_09:16:47]: iter 2650: avg_loss 0.008, avg_acc 0.600, step_time 4751.0ms, mfu 105.2%, throughput 88k tok/s, total time 16475s [2024-05-16_09:18:46]: iter 2675: avg_loss 0.008, avg_acc 0.680, step_time 4762.0ms, mfu 104.9%, throughput 88k tok/s, total time 16594s [2024-05-16_09:20:46]: iter 2700: avg_loss 0.014, avg_acc 0.520, step_time 4756.4ms, mfu 105.1%, throughput 88k tok/s, total time 16714s [2024-05-16_09:22:46]: iter 2725: avg_loss 0.008, avg_acc 0.640, step_time 4740.9ms, mfu 105.4%, throughput 88k tok/s, total time 16834s [2024-05-16_09:24:46]: iter 2750: avg_loss 0.008, avg_acc 0.640, step_time 4753.8ms, mfu 105.1%, throughput 88k tok/s, total time 16954s [2024-05-16_09:26:47]: iter 2775: avg_loss 0.009, avg_acc 0.560, step_time 4757.8ms, mfu 105.0%, throughput 88k tok/s, total time 17075s [2024-05-16_09:28:47]: iter 2800: avg_loss 0.008, avg_acc 0.640, step_time 4755.7ms, mfu 105.1%, throughput 88k tok/s, total time 17195s [2024-05-16_09:30:47]: iter 2825: avg_loss 0.008, avg_acc 0.520, step_time 4750.9ms, mfu 105.2%, throughput 88k tok/s, total time 17315s [2024-05-16_09:32:47]: iter 2850: avg_loss 0.008, avg_acc 0.640, step_time 4752.0ms, mfu 105.2%, throughput 88k tok/s, total time 17435s [2024-05-16_09:34:47]: iter 2875: avg_loss 0.008, avg_acc 0.680, step_time 4730.1ms, mfu 105.7%, throughput 88k tok/s, total time 17555s [2024-05-16_09:36:47]: iter 2900: avg_loss 0.009, avg_acc 0.680, step_time 4742.2ms, mfu 105.4%, throughput 88k tok/s, total time 17675s [2024-05-16_09:38:47]: iter 2925: avg_loss 0.005, avg_acc 0.720, step_time 4925.5ms, mfu 101.5%, throughput 85k tok/s, total time 17795s [2024-05-16_09:40:46]: iter 2950: avg_loss 0.007, avg_acc 0.720, step_time 4741.7ms, mfu 105.4%, throughput 88k tok/s, total time 17914s [2024-05-16_09:42:46]: iter 2975: avg_loss 0.007, avg_acc 0.760, step_time 4719.0ms, mfu 105.9%, throughput 89k tok/s, total time 18034s [2024-05-16_09:45:45]: loss estimation took 69.2 seconds. (2.3% of loop) [2024-05-16_09:45:45]: step 2999: train loss 4.4303, val loss 4.4114 [2024-05-16_09:47:12]: saving checkpoint to /app/suno/checkpoints/2024-05-16_04-33-58 [2024-05-16_09:54:25]: saving took 520.2 seconds. (17.4% of loop) [2024-05-16_09:54:37]: iter 2999: avg_loss 0.007, avg_acc 0.708, step_time 600720.4ms, mfu 0.8%, throughput 1k tok/s, total time 18745s [2024-05-16_09:54:37]: done. compute-hpc-node-706:1193446:1223374 [0] init.cc:1909 NCCL WARN Cuda failure 'out of memory' compute-hpc-node-706:1193446:1223374 [0] init.cc:2030 NCCL WARN commReclaim: comm 0x8116860 (rank = 18) in abort, error 1 compute-hpc-node-627:257681:287647 [0] init.cc:1909 NCCL WARN Cuda failure 'out of memory' compute-hpc-node-627:257681:287647 [0] init.cc:2030 NCCL WARN commReclaim: comm 0x7f13200 (rank = 13) in abort, error 1 compute-hpc-node-284:2851318:2881694 [0] init.cc:1909 NCCL WARN Cuda failure 'out of memory' compute-hpc-node-284:2851318:2881694 [0] init.cc:2030 NCCL WARN commReclaim: comm 0x917e770 (rank = 4) in abort, error 1