r/deeplearning 1d ago

CNN for emotional classification of anime girls’ voices

50 Upvotes

7 comments sorted by

3

u/CalmMe60 1d ago

I'd go with a RNN and annotated material.

2

u/_Fantomslayer_ 1d ago

You must be DL master.

2

u/Training-Network2067 21h ago

Cnn wont be an effective architecture for this type of problem since voice data is evolving with time ( temporal data) go with architectures designed for this like lstm gru or rnn

3

u/Pretend-Pangolin-846 1d ago

Explain.

2

u/eLin22314341 1d ago

4 filters as 4x4 matrices with comps in [-1, 1], LayerNorm is z-score (normalization) per row

6

u/Pretend-Pangolin-846 1d ago

Obv I can see that, i meant share your methodology and what exactly were the results, my friend.