Showing posts with label conference recap. Show all posts
Showing posts with label conference recap. Show all posts

Monday, October 13, 2014

Interspeech 2014 Recap

This year's Interspeech was in Singapore.    Singapore is, in some ways, a very easy venue to travel to.  It's a modern, cosmopolitan city.  They speak English.  It's tropical, but you're never more than a hundred meters from air conditioning.   In other ways, it's so far.  Over 20 hours each way.  I like airplanes.  They're as magical as any technology we've got.  But 20 hours is a long time to sit still.  Think about how many steps you take in 20 hours.  How many different faces you see.  Then reduce that to about 200 steps, and 10 people.

"Because we are all poets or babies in the middle of the night, struggling with being." - Martin Amis "London Fields"

Interspeech 2014 was a well run conference.  The quality of papers was generally quite high. The venue easily handled the size of the event and the wifi was steady.  It was difficult to find enough food at the Welcome Reception, but easy to find enough beer.  The banquet was flawed -- segregating vegetarians is pretty rude -- but they all are, and at least there was enough food to go around, and everyone ate promptly.  And of course, there was Mambo and Jambo.  I'm not going to go into it here, but find someone who attended the opening ceremony and ask them to describe it.  Then don't believe them, and ask someone else to do the same.  It was "odd" at best.  But be sure, I'll be attending the Dresden opening ceremony to see how they one-up it.

Deep Learning
A few years ago, DNNs invaded speech conferences.  Deep Learning is still a significant buzz word, and a hot topic.  But for, the better of everyone involved, the intensity has cooled.  Now the interest in DNNs seems to have shifted into 1) understanding how they work, and how to best train them to a task and 2) Long Short-Term Memory units to model sequential data.  The latter really broke out at this years conference.  There were a number of papers finding that them to be an effective alternative to traditional recurrent nets trained with back-propagation through time.

BABEL
I've been involved with the IARPA-BABEL program, so my view is pretty biased on this front, but I felt like the presence of BABEL in this year's Interspeech was particularly large.  The program's central task is performing keyword search on low-resource languages. It has an aggressive evaluation schedule with an increased number of *new* languages involved each year.  There were at least two sessions devoted to keyword search, and papers evaluated by either BABEL-proper or the NIST OpenKWS challenge seemed to be all over the conference.  (Searching the paper index suggests that there are between 30 and 40 BABEL papers, and another 10 or so OpenKWS papers.)  It seems clear that this program has had a large impact on ASR and KWS research.  2014 was certainly the high-water mark here, as the program shrunk by 50% last year, but it's worth noting its effect on the field.

Some standout papers

I don't mean to suggest that these are "the best" papers, but they're ones that caught my eye for one reason or another.

Acoustic Modeling with Deep Neural Networks Using Raw Time Signal for LVCSR by Zoltán Tüske, Pavel Golik, Ralf Schlüter and Hermann Ney.   Part of the promise of Deep Learning is the ability to "learn feature representations" directly from data.  This is frequently touted as a description of what is happening in the first hidden layer of a deep net.  So, the logic goes, do we need MFCC/PLP/etc. features, or can we do speech recognition directly on a raw acoustic signal?  This is the first paper I'm aware of to affirmatively show that "yes, yes we can".  It requires a good amount of training data, and rectified linear (ReLU) neurons work better for this, but 1) it works competitively with traditional features and 2) many of the first hidden layer neurons can be shown to be learning a filterbank.  Very cool.

Backoff Inspired Features for Maximum Entropy Language Models by Fadi Biadsy, Keith Hall, Pedro Moreno and Brian Roark.  In n-gram language modeling, when a sequence of words A, B, C have never been observed, its n-gram probability P(C|A,B) can be approximated by the probability P(C|B).  But to be sensible about it, you've got to apply a backoff penalty.  Discriminative language models can seamlessly incorporate backoff features, F(A,B,C), F(B,C), etc. and learn appropriate weights.  The key insight in this paper is that when a discriminative model uses these backoff estimates, it incurs no penalty.  It's essentially overestimating the probability of uncommon n-grams in the context of common (n-k)-grams.  This paper seeks to fix this, and does.

Additional favorites some from my students:

  • Word Embeddings for Speech Recognition by Samy Bengio and Georg Heigold.  Far and away the most popular poster at the conference.  This is high on my list to read closely.  It promises to learn a euclidean space into which word decoding can happen, so that words that sound similar are closer in space.  
  • Learning Small-Size DNN with Output-Distribution-Based Criteria by Jinyu Li, Rui Zhao, Jui-Ting Huang, Yifan Gong.  How do you effectively train a small DNN without ruining its performance?  This paper out of MSR suggests training a large DNN, then using its output distribution to train the small one.  It'll take me some closer reading to fully understand why this works, but I'm intrigued.





Wednesday, June 05, 2013

Deep Thoughts on ICASSP 2013

ICASSP 2013 is wrapping up today in Vancouver.  Unfortunately, I missed the last day (and sessions on speech synthesis and prosody that I would have enjoyed).  But a wedding on Saturday brought me back a day early.

I hadn't been to ICASSP before, mostly due to timing oddities and writing grants over summers rather than writing papers that would hit the deadline.  It is a very large conference.  About twice as large as Interspeech.  But the scope is also much broader.  Speech and Language work made up at most 30% of the work at the conference.  And even this is generous, including machine learning, and other work on audio.

So take this recap with a grain of salt.  I missed the last day of the conference, and my impressions are speech focused.  (I think I've described all conference recaps as blind-men-and-the-elephant problems and this one is no exception.)

Deep Learning.
OK, I pointed out that Deep Neural Nets were a "hot topic" at last years Interspeech.  It's hard to believe it's possible, but they're even hotter now.  Geoffrey Hinton gave the first plenary talk.  This was followed by an oral session called "Automatic Speech Recognition using Neural Networks", which was followed by a Special Session titled "New Types of Deep Neural Network Learning for Speech Recognition and Related Applications".  The next morning, you could attend "Acoustic Modeling with Neural Networks".  And this is just at the session level.  Even more applications of multilayer neural networks were scattered around other oral and poster sessions.  Some of these oral sessions were so crowded that people were standing along the walls and sitting in the aisles.  Nothing else that I saw received nearly so much attention.

It's easy to view "deep" learning as a silver bullet -- the next great machine learning that will solve all of our problems.  It's almost certainly not.  However, a wide array of research groups are seeing similar impressive performance gains by using deep network models for a broad spectrum of spoken language processing tasks.  This is especially true for acoustic modeling in speech recognition.  Given this, deep learning shouldn't be ignored.

Hinton's coursera course is a solid place to start. (Though resist drinking the kool-aid.  To my mind, perceptrons are bad approximations of neurons and worse approximations of the brain, and do little to advance our understanding of human intelligence.)

Another highlight
One paper which caught my attention for its simplicity came out of Google: "Language Model Verbalization for Automatic Speech Recognition".  Essentially "verbalization" is defined as a sort of inverse text-normalization. In text normalization for speech synthesis we have to translate "10" to "TEN", and "7:11" to "SEVEN ELEVEN" or "ELEVEN PAST SEVEN".  For ASR, the idea of verbalization is to convert decoding output of "SEVEN ELEVEN" into "7:11" or "7-11".  Why bother?  Well, Google (and everyone else) has big language models based on text data. You could run a text normalizer over all of this data, but the proposition here is to convert the ASR output into a form that looks more like the source material in your language model.

The Verbalizer solution to this problem is remarkably elegant.  A traditional WFST decoder can be expressed as D = C • L • G, where C comes off the acoustic model mapping context dependent to independent phones, L is the pronunciation model and G the language model.  The "Verbalized" WFST model includes a WFST V which maps ASR realizations like "SEVEN ELEVEN" to text-like realizations like "7-11" or "7:11" (and since it's a WFST it can do both simultaneously).  The new decoder looks like D = C • L • V • G.  No fuss, no muss.  Except that you have to write Verbalizer rules by hand.

The paper focused on terms involving numbers, but the framework is very extensible.  And it's great to see work coming out of Google that doesn't have Google-scale data as a prerequisite.

Meta-comment
The acceptance rate at this years ICASSP was 52%.  This means that the ICASSP and Interspeech acceptance rates are identical for the first time.  I know that Interspeech organizers have been working to lower the acceptance rate, while it sounds like there has been pressure to keep the size of ICASSP large, even at the expense of a higher acceptance rate.  IEEE (the ICASSP parent organization) is a much larger bureaucracy than ISCA (Interspeech).  There are clear expectations from IEEE about the expected revenue from hosting a conference, which translates to expectations on attendance and therefore the number of accepted papers regardless of the number of submissions.

Despite the near constant rain, I genuinely enjoyed Vancouver and ICASSP 2013.  I'm looking forward to the next.



Monday, September 17, 2012

Interspeech 2012 Recap

Portland proved to be a great venue for this year's Interspeech.  (Though people who attended ACL 2011 probably already could have guessed that.)

Setting up three simultaneous poster sessions in the parking garage may not sound like the mark of a good conference, but it was perfect.  There was loads of space between each presenter.  It allowed for all three sessions to be in the same place.  And the folks at the Hilton did a great job of making it fairly unrecognizable as a parking lot.  (In fact, Alejna Brugos didn't realize it until they were removing the carpets and "walls" on Thursday afternoon.)

Deep Neural Networks.
For "trends", there's really nothing hotter right now than Deep Neural Networks or Deep Belief Nets.  This isn't an area that I do research in, but the story goes more or less this.  Neural Networks with more than a few hidden layers don't train very well with back-propagation. Geoff Hinton and his group figured out how to overcome this limitation not too long ago.  (I think this 2006 paper explains it, but I can't be 100% sure.) Then at ASRU 2012 and ICASSP 2011 and 2012, the folks at Microsoft showed that you can use Deep Neural Networks to generate *very* useful front end features.  (Tara Sainath has a nice recap of ICASSP 2012 here.) Now, everyone wants a piece.

The field has expanded from Microsoft to include IBM and Stanford/Berkeley/Google and RWTH Aachen.  Joining them with posters on Deep Neural Nets for ASR are Tsinghua, CMU, Karlsruhe, NTT, INESC-ID, UWashington, and Georgia Tech.  At this point, there's no way to deny that this approach is receiving significant research attention.  The results seem to be holding up.  If only they didn't take so long to train...

Prominence Special Session.
I was particularly looking forward to the Special Session on Prominence.  On balance I was happy about the session.  It attracted work and discussion of prosody in a way that can sometimes feel diffuse and unfocused at a large conference like Interspeech.

I found this session to be surprising in a few ways.

It's been my understanding that "prominence" was used as a catch-all term to cover diverse prosodic phenomena including stress, emphasis, and pitch accenting.  The first surprising element of this Session came in a review of the paper I submitted to it.  The paper is on the use of automatically predicted pitch accents and intonational phrase boundaries to improve pronunciation modeling.  The review, while generally positive, found the paper to not be appropriate for a prominence session because it explored the use of "pitch accents" rather than "prominence".  I still haven't gotten a good explanation of the difference, and the reviews are blind.

A second surprise is that there seems to be a movement away from a phonological theory of prosody. Mark Hasegawa-Johnson and Jennifer Cole have been doing work over the last few years investigating how naive listeners perceive prominence.  They've consistently found that listeners respond to different qualities sometimes at different thresholds when assessing prominence.  I've found this line of research to be interesting and generally informative, but not a clear indictment of the theory that there perceptual and productive prosodic categories exist.    The panel (which I was a part of) on balance seemed comfortable with the idea that prominence is a continuous rather than categorical phenomenon.  This view was most directly expressed Denis Arnold who said approximately: focus can be categorical, stress can be categorical, while prominence is still continuous. I didn't understand this statement then, and still don't.  But again, this may be due to a different definition of prominence than I use.

The last surprise comes from finding out that there is a direction of pursuing language universals in prominence and prosody more broadly. Petra Wagner and Fabio Tamburini (the session organizers) are planning a workshop to investigate this.  In my experience, while the dimensions of prosodic variation may be used in multiple languages and some of these (e.g. increased intensity or duration) may be used to indicate prominence in all languages, it is extremely unlikely that either the communicative impacts of prosodic variation or its realization and perception are language-universal.   From that perspective, I'm not quite clear about what this line of research hopes to accomplish, but I'm curious about where it ends up.

Dynamic Decoding.
It appears that every year, I find myself sitting in on an oral session on a topic that I know very little about.  Last year it was the language identification session.  This year it was Dynamic Decoding.  I was most intrigued by this because I hadn't heard the term before.  When I asked someone what it was, they said "I don't know, Viterbi?".

I'm not quite sure this is a good enough distillation of the topic, but the papers in this session were about how to make on-the-fly (or post-training) modifications to language or pronunciation models.  This is a cool idea with clear practical importance -- how do you add words to a recognizer on a mobile device and have this appropriately incorporated into the LM and pronunciation model?  These two papers have some interesting WFST based approaches on this task.  I'll be curious to see learn more about this.  Also, if anyone has a more precise definition of this research area, I'd love to hear it.

Finally, some comments on two of the keynotes.

There were four keynotes at this year's Interspeech, two were about interesting inter/multi-disciplinary questions about how speech processing intersects with music and animal vocalization, respectively.


Chin-Hui Lee: An Information-Extraction Approach to Speech Analysis and Processing
A third was delivered by this years ISCA medalist, Chin-Hui Lee.  Prof. Lee's most famous accomplishment is MAP adaptation in acoustic modeling.  This is a researcher who spent a career treating speech recognition as a pattern matching problem.  This is a view embodied by the Fred Jelenik quote: "Every time I fire a linguist, the performance of the speech recognizer goes up".  What struck me, is that despite this view, in a talk summarizing a successful career, Prof. Lee presented a view of speech recognition that says that linguistic knowledge and speech science should be incorporated into the task.  This is an alternate perspective that has been investigated by a lot of talented researchers, including Hynek Hermansky, Jennifer Cole, Mark Hasegawa-Johnson, Alex Waibel, Hermann Ney (via speech-to-speech translation), Mari Ostendorf, Elizabeth Shriberg, Andreas Stolcke, Rene Beutler, Karen Livescu and many more (my apologies to anyone I missed).

I was struck by the evolution of perspective from someone who represents the statistical pattern matching approach to recognizing the potential importance of linguistic knowledge.


However, this talk was not so well received by some members of the audience for fairly obvious reasons.  Firstly, it over-played the importance of Prof. Lee's own contributions.  In a slide on "My contributions", virtually all major improvements to ASR over the last 20 years were mentioned including most styles of adaptation (including MAP), and virtually all major forms of discriminative training.  Secondly, it failed to recognize that the linguistic inspired approach that he was advocating for the future had been extensively researched by other talented peers.


On balance, I found it a compelling message.  In principle, it understandably rubbed some people the wrong way.


Michael Riley: Weighted Transducers in Speech and Language Processing
I should preface my comments about Michael Riley's keynote by saying that we worked together while I was interning at Google.  I'm a fan.  Michael has the rare quality of being the smartest guy in the room without letting anyone know until its genuinely useful.

The best part of this keynote was the history of the Weighted Finite State Transducer.  This was a great story that takes place largely at Bell Labs in the 90s and features Fernando Pereira, Mehryar Mohri and, naturally, Michael Riley.  This section was appropriately personal, while presenting this relevant recent history.  The WFST is so ubiquitous in speech and NLP applications that it's easy to forget that it's has a human context.

Much of the rest of the keynote felt like a 3 hour tutorial compressed into 40 minutes.  This involved showing algorithms, and example WFSTs and describing all of the things that they can be used for.  While a successful demonstration of the breadth of application, it was presented at such a pace that it was difficult to get anything out of it, if you didn't know it already.  I'd point the interested to the references found on the OpenFST page for more thorough tutorials that can be digested at your own pace.

Interspeech 2012 was successful and fun.  Portland and the Hilton (and it's solid wifi) were excellent hosts.  There was good work and as ever more than I could see.  If you have great or favorite papers that I missed, please let me know!



Tuesday, September 06, 2011

Interspeech 2011 Recap

Interspeech 2011 was held in Florence Italy about a week ago, August 28-August 31.  A vacation in Italy was too good to pass up on, so The Lady joined me, and we stayed until Labor Day.

I ended up sending a bulk of work to Interspeech, so spent more time than usual in sessions that I was presenting in rather than seeing a lot of papers.

Two interesting themes stood out to me this year.  Not for nothing, but these represent some novel ideas about speech science through understanding dialog and the speech engineering.

Entrainment
Julia Hirschberg, my former advisor, received the ISCA medal for her years of work in speech.  Her talk was on current work with Agustín Gravano and Ani Nenkova on entrainment.  Entrainment is the phenomenon by which when people are speaking to each other, their speech becomes more similar.  This can be realized in terms of the words that are used to describe a concept, as well as speaking rate, pitch, intensity.  How this happens isn't totally understood, and measures of entrainment are still being developed.  This research theme is still in its early phases, but I haven't seen an idea spread around a conference as quickly or as thoroughly as this did.  There were questions and discussions all over the place (like Tom Mitchell's keynote about fMRI data and word meaning) about this phenomenon.  The more engineering folks weren't as compelled by the utility of this in helping speech processing, but within the speech perception and production communities, and specifically the dialog and IVR folks, it was all the rage.  It'll be something to see how this develops.


I-vectors
I-vectors are a topic that I need to spend some time learning.  I hadn't heard of it before this conference, where there were no less than a dozen papers that used this approach.  Essentially the idea is this:  The location of mixture components in a GMM model are composed of (in at least one form) a UBM, a channel component, and a "interesting" component.  This "interesting" component can be the contribution of a particular speaker, or a language/dialect, or anything else you're trying to model.  Joint Factor Analysis is used to decompose the observation into these components in an unsupervised fashion.  It's in this part where my understanding of the math is still limited.  The crux is that the "interesting" component can be represented by a \[Dx\] transformation, where the dimensionality of x can be set by the user.  In comparison to a supervector representation, where the dimensionality of the supervector is constrained to be equal to the number of parameters (or means) of the GMM, i-vectors can be significantly smaller leading to better estimation and smaller models.  I'll be reading this tutorial by Howard Lei over the next few weeks to get up to speed.


There were few specific papers that stood out to me this conference.  I'm intrigued by Functional Data Analysis as a way to model continuous time-value observations. Michele Gubian gave a tutorial on this that I sadly missed, and included it in at least one paper, Predicting Taiwan Mandarin tone shapes from their duration by Chierh Chung and Michele Gubian.  This paper wasn't totally convincing in the utility of the technique, but there may be more appropriate applications.


It was a satisfying and inspiring conference, to be sure.  I think I was more interested in talking to people than in papers in particular this time.  If anyone has particular favorites that I missed, please use the comments to share or just email me.


Tuesday, October 05, 2010

Interspeech 2010 Recap

Interspeech was in Makuhari, Japan last week.  Makuhari is about 40 minutes from Tokyo, and I'd say totally worth the commute.  The conference center was large and clean, and (after the first day) had functional wireless, but Makuhari offers a lot less than Tokyo does.

Interspeech is probably the speech conference with the broadest scope and largest draw.  This makes it a great place to learn what is going on in the field.

One of the things that was most striking about the work at Interspeech 2010 was the lack of a Hot Topic.  Acoustic modeling for automatic speech recognition is a mainstay of any speech conference, that was there in spades.   There was some nice work on prosody analysis.  Recognition of age, affect and gender were highlighted in the INTERSPEECH 2010 Paralinguistics Challenge, but outside the special session focussing on this, there wasn't an exceptional amount of work on these questions.  Despite the lack of a new major theme to emerge this year, there was some very high quality, interesting work.

Here is some of the work that I found particularly compelling.
  • Married Couples' speech
    Sri Narayanan's group with other collaborators from USC and UCLA have collected a set of married couples' dialog speech during couple's therapy. So this is already compelling data to look at.  You've got naturally occurring emotional speech, which is a rare occurrence, and it's emotion in dialog.  They had (at least) 2 papers on this data at the conference, one looking at prosodic entrainment during these dialogs, and the other classifying qualities like blame, acceptance, and humor in either souse.  Both very compelling first looks at this data.  There are obviously some serious privacy issues with sharing this data, but hopefully it will be possible eventually.

    Automatic Classification of Married Couples’ Behavior using Audio Features Matthew Black, et al.

    Quantification of Prosodic Entrainment in Affective Spontaneous Spoken Interactions of Married Couples
    Chi-Chun Lee, et al.
  • Ferret auditory cortex data for phone recognition
    Hynek Hermansky and colleagues have done a lot of compelling work on phone recognition.  To my eye, a lot of it has been banging away at techniques other than MFCC representations for speech recognition.  Some of them work better than others, obviously, but it's great to see that kind of scientific creativity applied to a core task for speech recognition.  This time the idea was to take "spectro temporal receptive fields" empirically observed from ferrets that have been trained to be accurate phone recognizers, and use these responses to train a phone classifier.  Yeah, that's right.  They used ferret neural activity to try to recognize human speech.  Way out there.  If that weren't compelling enough, the results are good!
  • Prosodic Timing Analysis for Articulatory Re-Synthesis Using a Bank of Resonators with an Adaptive Oscillator Michael C. Brady
    A pet project has been to find a nice way to process rhythm in speech for prosodic analysis.  Most people use some statistic based on the intervocalic intervals, but this is unsatisfying.  While is captures the degree of stability of the speaking rate, it doesn't tell you anything about which syllables are evenly spaced, and which are anomalous.  This paper uses an adaptive oscillator to find the frequency that best describes the speech data.  One of the nicest results (that Michael didn't necessarily highlight) was that he found that deaccented words in his example utterance were not "on the beat".  In the near term I'm planning on replicating this approach for analyzing phrasing, on the idea that in addition to other acoustic resets, the prosodic timing resets at phrase boundaries. A very cool approach.
  • Compressive Sensing
    There was a special session on compressive sensing that was populated mostly by IBM speech team folks.  I hadn't heard of compressive sensing before this conference, and it's always nice to learn a new technique.  At its core compressive sensing is an exemplar based learning algorithm.  Where it gets clever is that where k-means uses a fixed number, k, of exemplars to use with equal weight, and SVM use a fixed set of support vectors to make decisions, in compressive sensing a dynamic set of exemplars are used to classify each data point.  The set of candidate exemplars (possibly the whole training data set) are then weighted with some L1-ish regularization to drive most of the weights to zero -- selecting a subset of all candidates for classification.  Then a weighted k-means is performed using the selected exemplars and weights.  The dynamic selection and weighting of exemplars outperforms vanilla SVMs, but the process is fairly computationally expensive.
Interspeech 2011 is in Florence, and I can't wait -- now I've just got to get banging on some papers for it.

Tuesday, June 08, 2010

HLT-NAACL 2010 Recap

I was at HLT-NAACL in Los Angeles last week.  HLT isn't always a perfect fit for someone sitting towards the speech end of the Human Language Technologies spectrum.  Every year, it seems, the organizers try (or claim to try) to attract more speech and spoken language processing work.  It hasn't quite caught on yet and the conference tends to be dominated by Machine Translation and Parsing.  However...The (median) quality of the work is quite high. This year I kept pretty close to the Machine Learning sessions and got turned on to the wealth of unsupervised structured learning which I've overlooked over the last N>5 years.


There were two new trends that I found particularly compelling this year:
  • Noisy Genre
    This is pretty clunky term covers genres of language which are not well-formed.  As far as I can tell this covers everything other than Newswire, Broadcast news, and read speech. This is what I would call "language in the wild" or in a snarkier mood, "language" (sans modifier). For the purposes of HLT-NAACL, it covers Twitter messages, email, forum comments, and ... speech recognition output.  It's this kind of language that got me into NLP and why I ended up working on speech, so I'm pretty excited that this is receiving more attention from the NLP community at large. 
  • Mechanical Turk for language tasks
    Like the excitement over wikipedia a few years ago, NLP folks have fallen in love with Amazon's Mechanical Turk. Mechanical turk was used for speech transcription, sentence compression, paraphrasing, and quite a lot more; there was even a workshop day dedicated solely to this topic. I didn't go to it, but will catch up on the papers this week or so.  This work is very cool, particularly when it comes to automatically detecting and dealing with outlier annotations.  The resource and corpora development uses of Mechanical Turk are obvious and valuable. It's in the development of "high confindence" or "gold standard" resources that I think this work has an opportunity to intersect very nicely in work on ensemble techniques and classifier combination/fusion.  If each turker is considered to be an annotator, the task of identifying a gold standard corpus is identical to generating a high-confidence prediction from an ensemble.
I had a sense of HLT-NAACL that was unfair:  My impression was that the quality of work was fairly modest.  I attribute this to three factors. 1) In the past there has been a lot of work of the type -- "I found this data set. I used this off-the-shelf ML algorithm. I got these results".  There's nothing particularly wrong with this type of work, except for it's boring, and not intellectually rigorous, it's not scientifically creative, and it doesn't illuminate the task with any particular clarity. (Ok, so there're at least four things wrong with this kind of work.)  2) HLT-NAACL accepts 4 page short papers.  With its formatting guidelines, it is almost impossible to fit more than a single idea in a 4 page ACL paper.  This leads to a good amount of simple or undeveloped ideas. (I've written a fair amount of these 4 page papers because they are accepted at a later deadline, but it's always frustrating when you realize you have more to say.)  3) And I think this is probably the most significant -- I've had a good amount of luck getting papers accepted to HLT-NAACL including my first publication in my first year of grad school.  This is probably just "I don't want to belong to any club that will accept people like me as a member"-syndrome, but it left me underestimating the caliber of this conference.

A couple of specific highlights of papers I liked this year:

  • “cba to check the spelling”: Investigating Parser Performance on Discussion Forum Posts Jennifer Foster.  This might be the first time I fully agree with a best paper award.  This paper looked at parsing outrageously sloppy forum comments. These are rife with spelling errors, grammatical errors, weird exclamations (lol).  The paper is a really nice example of the difficulty that "noisy genres" of text pose to traditional (i.e., trained on WSJ text) models.  The error analysis is clear and the paper proposes some nice solutions to bridge this gap by adding noise to the WSJ data. Also, bonus points for subtly including 
  • Cheap, Fast and Good Enough: Automatic Speech Recognition with Non-Expert Transcription
    Scott Novotney and Chris Callison-Burch.  A nice example of using Mechanical Turk to generate training data for a speech recognizer.  High quality transcription of speech is pretty expensive and critically important to speech recognizer performance.  Novotney and Callison-Burch found that Turkers are able to transcribe speech fairly well, and at a fraction of the cost.  This paper includes a really nice evaluation of Turker performance and some interesting approaches to ranking Turker performance.
  • The Simple Truth about Dependency and Phrase Structure Representations: An Opinion Piece
    Owen Rambow.  This paper was probably my favorite in terms of bringing joy and being a breath of fresh air. The argument Rambow lays out is that Dependency and Phrase Structure Representations of syntax are meaningless in isolation.  Moreover, these are simply alternate representations of identical syntactic phenomena.  Linguists love to fight over a "correct" representation of syntax.  This paper takes the position that the distinction between the representations is merely preference not substantive -- fighting over the correct representation of a phenomenon is a distraction to understanding the phenomenon itself.  Full disclosure: I've known Owen for years, and like him personally as well as his work.
  • Type-Based MCMC
    Percy Liang, Michael I. Jordan and Dan Klein.  Over the last few years, I've been boning up on MCMC methods.  I haven't applied them to my own work yet, but it's really only a matter of time.  This work does a nice job of pointing out a limitation of token based MCMC -- specifically that sampling on a token by token basis can make it overly difficult to get out of local minima.  Some of this difficulty can be overcome by sampling based on types, that is, sampling based on a higher level feature across the whole data set, as opposed to within 
    a particular token.  This makes intuitive sense and was empirically well motivated.

As a side note, I'd like to thank all you wonderful machine learning folks who have been doing a remarkable amount of unsupervised structured learning that I should have been paying better attention to over the last few years.  Now I've got to hit the books.

Monday, May 17, 2010

Speech Prosody 2010 Recap

Speech Prosody is a biannual conference held by a special interest group of ISCA.  Despite working on intonation and prosody since about 2006, this is the first year I've attended.

The conference has only been held five times and has the feeling of a workshop -- no parallel sessions is nice, but the quality of work was very varied.  I'm sympathetic to conference organizers and particularly sympathetic when you're coordinating a conference that is still relatively new.  I'm sure you want to make sure people attend, and the easiest way to do that is to accept work for presentation.  But by casting a wider net some work gets in that probably could have used another round of revision.  Mark Hasegawa-Johnson and his team logistically executed a very successful conference.  But (at the risk of having my own work rejected) I think the conference is mature enough that Speech Prosody 2012 could stand more wheat, less chaff.

Two major themes stuck out:

  • Recognizing emotion from speech -- A lot of work, but very little novelty in corpus, machine learning approach, or findings.  It's easy to recognize high vs. low activation/arousal, and hard to recognize high vs. low valence.  (I'll return to this theme in a post on Interspeech reviewing...)
  • Analyzing non-native intonation -- There were scads of papers on this topic covering both how intonation is perceived and produced by non-native speakers of a language (including two by me!). I had no anticipation that this would be so popular, but if this conference were any indication, I would expect to see a lot of work on non-native spoken language processing at Interspeech and SLT.

Here are some of my favorite papers. This list is *heavily* biased towards work I might cite in the future more than the "best" papers at the conference.

  • The effect of global F0 contour shape on the perception of tonal timing contrasts in American English intonation Jonathan Barnes, Nanette Veilleux, Alejna Brugos and Stefanie Shattuck-Huffnagel  The paper examines the question of peak timing in the perception of L+H* vs. L*+H accents.  Generally has been assumed (including in the ToBI guidelines) that the difference is in the peak timing relative to the accented vowel or syllable onset -- with L*+H having later peaks. Barnes et al. found that this can be manipulated such that if you keep the peak timing the same, but make the F0 approach to that peak more "dome-y" or "scoop-y", subjects will perceive a difference in accent type.  This is interesting to me, though *very* inside baseball.  What I liked most about the paper and presentation was the use of Tonal Center of Gravity or ToCG to identify a time associated with the greatest mass of f0 information.  This aggregation is more robust to spurious f0 estimates than a simple maximum, and is a really interesting way to parameterize an acoustic contour.  Oh, and it's incredibly simple to calculate: \[ T_{cog} = \frac{\sum_i f_i t_i}{\sum_i f_i}\]. 
  • Cross-genre training for automatic prosody classification Anna Margolis, Mari Ostendorf, Karen Livescu Mari Ostendorf and her students have been looking at automatic analysis of prosody for as long as anyone.  In this work, her student Anna is looking at how detection of accents and breaks performance degrades across corpora, looking at the BURNC (radio news speech) and Switchboard (spontaneous telephone speech) data.  This paper found that you suffer a significant degradation of performance when models are trained on one corpus and evaluated on the other, compared to within-genre training -- with lexical features suffering more dramatically than acoustic features (unsurprisingly...).  However, the interesting effect is that by using a combined training set from both genres, the performance doesn't suffer very much.  This is encouraging for the notion that when analyzing prosody, we can include diverse genres of training data and increase robustness.  It's not so much that this is a surprising result as it is comforting.  
  • C-PROM: An Annotated Corpus for French Prominence Study M. Avanzi, A.C. Simon, J.-P. Goldman & A. Auchlin This was presented at the Prominence Workshop preceding the conference.  The title pretty much says it all. But it deserves acknowledgement for sharing data with the community.  Everyone working with prosody bemoans the lack of intonationally annotated data.  It's time consuming, exhausting work. (Until last year I was a graduate student -- I've done more than my share of ToBI labeling, and I'm sure there will be more to come.)  These lovely folks from France (mostly) have labeled about an hour of data for prominence from 7 different genres, and ... wait for it ... posted it online for all the world to see, download and abuse.  They've set a great example, and even though I haven't worked with French before and have no intuitions about it, I appreciate open data a lot.
Full disclosure: I had two papers at Speech Prosody 2010, one in the main workshop (circuitously placed in a Special Session) and the other in a workshop on prominence the day before.  And I would certainly include these in the "varied" description of the caliber of work.  Not that I'd describe my work as shoddy papers (though there were some at SP2010, to be sure) but they are quite limited in scope, reporting on two small perception and production studies of non-native intonation.