# Annotating video and text simultaneously

**URL:** <https://community.labelstud.io/t/annotating-video-and-text-simultaneously/558>\
**Category:** Label Studio Support\
**Created:** [May 22, 2025, 7:39am UTC](https://community.labelstud.io/t/annotating-video-and-text-simultaneously/558 "2025-05-22T07:39:01Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Malo\_Maisonneuve](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.labelstud.io/malo_maisonneuve/32/342_2.png) [@Malo\_Maisonneuve](https://community.labelstud.io/u/Malo_Maisonneuve)\
**Post date:** [May 22, 2025, 7:39am UTC](https://community.labelstud.io/t/annotating-video-and-text-simultaneously/558/1 "2025-05-22T07:39:01Z")

</div>

Hello,

I’m currently trying to build a setup where we could simultaneously annotate data that is composed of synced text and video. The text corresponds to an automatic transcription of the audio signal from the video.  
Then, labels are used to classify movements occuring in the video, using transcribed text as details to support the labeling process.  
Thus, text should appear as one line, below the video, and should scroll automatically as the video is scrolling.

Is it something that can be done currently in the interface? And if not from initial object, can I create specific object to handle this scenario?

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![makseq](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.labelstud.io/makseq/32/26_2.png) [@makseq](https://community.labelstud.io/u/makseq)\
**Post date:** [May 22, 2025, 2:35pm UTC](https://community.labelstud.io/t/annotating-video-and-text-simultaneously/558/2 "2025-05-22T14:35:11Z")

</div>

Hi, try using this one if you need a single choice for the whole video:

```auto
<View>
  <Header value="Video Movement Classification with Synced Transcript" size="3" style="margin-bottom: 1em;"/>
  <Video name="video" value="$video" height="400" sync="audio" />
  <Audio name="audio" value="$video" sync="transcript" />
  <Paragraphs name="transcript" value="$transcript" layout="dialogue" 
              contextScroll="true" audioUrl="$video" sync="audio" />
  
  <Choices name="movement" toName="video" choice="single">
    <Choice value="Walking" />
    <Choice value="Running" />
    <Choice value="Jumping" />
    <Choice value="Sitting" />
    <Choice value="Other" />
  </Choices>
</View>

<!--{
  "video": "/static/samples/opossum_snow.mp4",
  "transcript": [
    {"author": "Speaker", "text": "The subject enters the frame and starts walking.", "start": 0, "end": 2},
    {"author": "Speaker", "text": "Now the subject is running quickly across the field.", "start": 2, "end": 5},
    {"author": "Speaker", "text": "The subject jumps over an obstacle.", "start": 5, "end": 7},
    {"author": "Speaker", "text": "Finally, the subject sits down to rest.", "start": 7, "end": 10}
  ]
}
-->

```

Here’s another option if you need labels across the timeline:

```auto
<View>
  <Header value="Video Movement Classification with Synced Transcript" size="3" style="margin-bottom: 1em;"/>
  <Video name="video" value="$video" height="400" sync="audio" />
    <TimelineLabels name="timelineLabels" toName="video">
    <Label value="Walking" />
    <Label value="Running" />
    <Label value="Jumping" />
    <Label value="Sitting" />
    <Label value="Other" />
  </TimelineLabels>
  
  <Audio name="audio" value="$video" sync="transcript" />
  <Paragraphs name="transcript" value="$transcript" layout="dialogue" 
              contextScroll="true" audioUrl="$video" sync="audio" />
  
</View>

<!--{
  "video": "/static/samples/opossum_snow.mp4",
  "transcript": [
    {"author": "Speaker", "text": "The subject enters the frame and starts walking.", "start": 0, "end": 2},
    {"author": "Speaker", "text": "Now the subject is running quickly across the field.", "start": 2, "end": 5},
    {"author": "Speaker", "text": "The subject jumps over an obstacle.", "start": 5, "end": 7},
    {"author": "Speaker", "text": "Finally, the subject sits down to rest.", "start": 7, "end": 10}
  ]
}
-->

```

---

<div class="post-metadata">

**Author:** ![Malo\_Maisonneuve](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.labelstud.io/malo_maisonneuve/32/342_2.png) [@Malo\_Maisonneuve](https://community.labelstud.io/u/Malo_Maisonneuve)\
**Post date:** [May 23, 2025, 9:22am UTC](https://community.labelstud.io/t/annotating-video-and-text-simultaneously/558/3 "2025-05-23T09:22:18Z")

</div>

Thank you very much for your answer!  
The second example you provided corresponds to my use case.  
The only issue I’m encountering now is that the text is not scrolling automatically, even with the option enabled. It is also showing the first text (“The subject enters the frame and starts walking.”) at the end of the scrollable paragraph area, even though its start time is 0.  
Am I missing some options?  
I’m using the Community Label Studio Version.

---

<div class="post-metadata">

**Author:** ![makseq](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.labelstud.io/makseq/32/26_2.png) [@makseq](https://community.labelstud.io/u/makseq)\
**Post date:** [May 23, 2025, 9:57am UTC](https://community.labelstud.io/t/annotating-video-and-text-simultaneously/558/4 "2025-05-23T09:57:19Z")

</div>

When you tried my exact example, did it work correctly?

---

<div class="post-metadata">

**Author:** ![Malo\_Maisonneuve](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.labelstud.io/malo_maisonneuve/32/342_2.png) [@Malo\_Maisonneuve](https://community.labelstud.io/u/Malo_Maisonneuve)\
**Post date:** [May 23, 2025, 10:52am UTC](https://community.labelstud.io/t/annotating-video-and-text-simultaneously/558/5 "2025-05-23T10:52:03Z")

</div>

Your example works perfectly.  
I narrowed down the problem to using a webm video file.  
After a conversion to mp4, the paragraph scrolls correctly through the texts.  
Thanks!
