|
|
|
|
|
by zhinit
12 days ago
|
|
the spectrograms are 128x173 (128 mel frequency bins by 173 time frames)
the encoder is downsampling 4 stages of stride 2 convolutions so it halves dimensions 4 times 0: 128 x 173 1: 64 x 87 2: 32 x 44 3: 16 x 22 4: 8 x 11 Then i used 4 separate channels. This was somewhat arbitrary due to the local training constraint. This would be a hyper parameter worth tuning if I had time to dig into this more. I trained this a few month ago and don't remember exactly what I tried before I arrived here, but I only ran the whole process 2 or 3 times because of how long it took to train.
Hope this answers your question! |
|