Showing posts with label listening. Show all posts
Showing posts with label listening. Show all posts

Speech Intelligibility and Grain Size - a SLRF 2017 preview post

Next month, I'll be giving a talk called "Explaining Intelligibility: What matters most in L2 Speech?" at Second Language Research Forum 2017 in Columbus, Ohio. That talk will examine features of L2 Korean speech that caused intelligibility issues, based on data from 30 Korean native speaking listeners. This post is a preview, where I'll show some of my initial summary data.

In second language speech, a handful of constructs are widely studied and considered important: intelligibility, comprehensibility, accentedness, and fluency. Arguably, intelligibility is the most important, as you can't really have successful communication without it. When we think about intelligibility, we might think of it holistically to describe a person's general ability or a person's performance in some speaking context. Language tests are a good example of this- the word "intelligibility" pops up in rubrics that are used to assess someone's speaking performance on a test task. Visually, we might think of people having different degrees of intelligibility looking like this:

Fig 1. Average proportion of eojeols (words) in a picture description task correctly transcribed by 30 Korean listeners. 
In Fig. 1, we do see some variation- around 80% or so of Speaker B's words (actually 어절, eojeol, a word + bound morphemes, the preferred unit of analysis in Korean linguistics) were intelligible to Korean listeners, on average, while speaker F clocked in at around 50%.  But it isn't necessarily the case that Speaker B is always 30% more intelligible than Speaker F- everybody stumbles sometimes, right? What if we look at each utterance (sentence) that the speakers produced?
Fig 2. Average proportion of eojeols correctly transcribed in each utterance.
We can see here that Speaker F, while generally having troubles with intelligibility, really dropped the ball on his/her first sentence, which was almost completely unintelligible to the listeners in the study. Speaker B is relatively intelligible throughout, but his/her first sentence was a little harder to grasp compared to the following two. Speaker A shows one of the starkest contrasts, with his/her first sentence around 50% and the rest being 80% or so. It's worth noting that the person and sentence level is about as fine-grained as a lot of L2 intelligibility research using naturalistic or contextualized speech has gone. Some studies focused on single-word intelligibility (i.e., a learner reads single words, or names single objects) do get to the word level, but I am curious about what leads to intelligibility issues in more realistic contexts. After all, these utterances aren't uniformly 50% intelligible- each word is either intelligible, or not. So we can dial in here and look at things this way:

Fig 3. Proportions of correct transcriptions for each eojeol.
To me, this is where things get really interesting. For one, we can see much more variation- there's more red and orange in this plot compared to the utterance-level depiction in Fig 2. Some words were almost completely unintelligible to listeners. Those who read Korean might notice that many of these words are names! This is interesting, and was intentional in the task design for the speakers- a name that involves a nasal assimilation at the meeting of its two syllables was chosen. But other words, often involving times and days of the week, were also quite difficult for listeners to understand. What I'm more interested in, though, is the speech features that might cause these words to be unintelligible- is it phoneme substitutions? Deletions? Pauses or repetitions in the utterance? Lexical errors? Grammatical errors? And that's my next task- building models to examine the relative impacts of these (and other) features on intelligibility.

Stay tuned!

P.S. - It's also worth pointing out that the 30 listeners were not monolithic in their overall ability to understand and correctly transcribe words. I'll be looking at listener factors in another analysis at a later time, but here's a little preview of that:

Fig 4. Proportion of eojeols correctly transcribed by each listener.

Using Qualtrics and Soundcloud to do speech/listening research online

Intro

The internet is making it easier and easier to conduct L2 research. It's really easy to send out questionnaires to language learners or teachers, for example. It's also a great tool for collecting writing samples, or having learners read short texts and answer questions. Importantly, the internet makes finding and interacting with participants much easier than doing everything in person. Although you don't get the same level of control over experimental conditions that are valuable for some kinds of research, I'd argue that the internet makes it much more feasible to get non-undergraduate participants at low or no cost. Essentially, you trade experimental control for better sampling (than you could get otherwise).

One thing that isn't really easy to do on the internet is responding to or recording audio. Getting audio samples from participants over the internet in response to some kind of stimulus is pretty much a nightmare, and one that I would love to see solved (honestly, it seems like internet-based test providers are the only ones with a solid handle on this, in terms of out-of-the-box solutions). However, it's becoming a bit easier to get audio samples to participants, and as long as they're responding with clicks or typing, we're in business.

In this post, I'll show you how I've used Qualtrics and Soundcloud to collect transcriptions for a speech intelligibility project I am currently working on. In this project, I had 30 native speakers of Korean (about half located in Korea, and the others spread around the globe) transcribe 28 utterances produced by L2 Korean learners.

Qualtrics

 Qualtrics is a well-known online survey platform. There's a good chance your institution has deluxe access to it that will allow you to use most available functions AND dress up your surveys with a nice, official-looking stylesheet. If not, you can still use Qualtrics for free as an individual; the free version will let you do many kinds of simple surveys.

For this post, I'm going to assume some basic familiarity with Qualtrics. If you're totally new to it, head over to support.qualtrics.com to read up on the basics (and honestly, it's pretty intuitive- just make an account and start playing around to get a feel for it).

Qualtrics does have built-in audio/video support, letting you upload files to your account and embedding them in survey questions. Going this route, you can very easily implement something like this:


 

To do this, create a new Text/Graphic question, type your direction (e.g., "Click play to listen.") and then click the Rich Content Editor... tab. Then, click the little film icon to upload and insert a media file.



After you have your audio file loaded in, create a new Text Entry question below- this is where a listener can type what they heard. You can also add any other type of question, or multiple questions. This is a nice, simple solution, but there are some potential problems. For one, the audio player is fully controllable by the participant. This means that a participant could listen multiple times, and that participants may all have different numbers of repeated listenings. This is a major trade-off in control! Along the same lines, participants can play/pause at will, and scrub back and forth. For some research tasks, this might be fine (say, if you're just farming out audio corpus transcription). But if you want to measure someone's listening comprehension or the intelligibility of a particular speaker/utterance, well, this doesn't give us enough experimental control to confidently do that.

Another weakness of the Qualtrics media player is that it uses Flash. Flash is a fairly common piece of web software (though waning in popularity and use). In my experience it's not quite ubiquitous and universal- every time you see that broken media link with a little puzzle piece on a webpage, it's because your version of Flash isn't up to date or is for some reason incompatible with what the page is trying to show you. As a researcher, you don't want to lose potential participants because they can't play your stimuli.

Embedded streaming audio with Soundcloud

One workaround is embedding streaming media from an external site. Soundcloud uses HTML5 to stream audio. HTML5 is a nearly universal standard on the web, and almost all contemporary browsers handle it well. And as I'll show you, you can customize how embedded Soundcloud audio displays on your Qualtrics survey to increase your level of experimental control.

Get yourself a Soundcloud account (you get free audio storage for about 2 to 3 hours worth of stuff), and upload a file. When you upload, make sure to set your file to Private if you don't want it to be accessible to just anyone and everyone on the internet. After uploading, click the Share button by your file and then click over to the Embed tab.



 If you want to be able to limit participants' control over audio playback, you'll want to click More Options near the bottom of the embed tab and click the checkbox for Enable automatic play. This might seem like a bad idea, but we'll build our own means of advancing through the survey that won't startle your participants with audio unexpectedly playing.


Finally, click in the Code and preview box to highlight your embed code. Hit ctrl+c to copy the code. Back to Qualtrics!









Over in Qualtrics, we're first going to make a question that allows participants to be ready to hear an audio stimulus. Create a new Descriptive Text question, and type a direction like "Click the >> button to play the next audio file." Next, insert a Page Break to require the participant to manually advance to the next question.

Now, create a new Text Entry question (or Multiple Choice, or whatever you'd like). Click on the HTML tab, and paste (ctrl-v) the embed code from Soundcloud. To get rid of any audio controls, change the value after iframe width to 0%, and change the value after height to 0. If you look carefully at the rest of the embed code, you can see that auto_play is set to true (you'll also notice some blurry stuff in the screenshot below- just keeping my private file private!). For the other options, you could go through and set them all to false, but this is ultimately unnecessary- since you made the Soundcloud embed a 0x0 pixel box, there's nothing that anyone can click on in your survey.



In final form, we get:
1. A screen that gets the participant ready to listen

 2. A screen where audio automatically starts playing, and a question that a participant can answer.


Closing 

I hope this post has been helpful. By inserting an automatically playing, one-time-only streaming audio file into Qualtrics, you can have a wide range of participants respond to audio stimuli while still maintaining control over the number of replays and play/pause functions. With some finesse, you could even add a limited number of chances for participants to replay audio (Qualtrics features some fairly robust logic and sequencing options).

Also, I strongly recommend including a practice item or two that serve as audio hardware checks so that participants can make sure their speakers/headphones are a) on, and b) at a comfortable listening volume. In my experiment, I had them transcribe a speaker saying "I can hear the audio well" as a technical check before going on to a practice item.

While you might not be able to achieve the level of control necessary for any speech or listening research, I think we're getting to the point where a lot can be accomplished with pre-packaged data collection software/platforms... if you're willing to do just a little tweaking.