Learning Joint Representation of Human Motion and Language

Contrastive learning for shared motion–language representations.

MoLang learns a shared representation of human motion and language from paired and unpaired data. The contrastive motion–language model supports action recognition and motion retrieval within a single representation space. (Kim et al., 2022)

Resources

PDF · arXiv

References

2022

  1. arXiv
    molang.png
    Learning Joint Representation of Human Motion and Language
    Jihoon Kim, Youngjae Yu, Seungyoun Shin, and 2 more authors
    arXiv preprint arXiv:2210.15187, Oct 2022