Sunday, April 27, 2014

Things I didn't know before becoming a professor: #1 How to Teach

In order to be a professor* you need a PhD.  The two responsibilities you have as a professor are to do good research and to teach well.

To get a PhD, you need to do good research, so you're well equipped to handle this.

To get my PhD, I did not have to teach.  I had TA'd a few times.  But I had put in 10 years in college/grad school.  I had taken countless classes, some with great teachers, some less great.  So I figured no big deal; I can teach.

I've seen good movies.  I've got a good story.  I can make a good movie.

These are some things I didn't know about teaching before becoming a professor.

You are television.
I learned this from Michael Cirino in a very different context.

For the time that you are in front of students, you are television.  You are putting on a show. If your students don't engage with you, your students won't learn anything beyond what they get in a book.     That's not to say that you're entertainment, but you are performing.

My act isn't strong yet, but it's getting better.

Don't skimp on the basics.
In an early version of a Machine Learning class, I decided that I wanted a lecture on spectral clustering. It was a topic I was interested in, but didn't have a ton of experience with. I read a lot. I worked out math.  By the time I had a lecture ready, I was feeling good about it.  I went to class, delivered my A material....blank stares.  It was a total dud.

It wasn't that the lecture was bad, but I hadn't earned it.  Because I wanted to get to a topic that was exciting to me, I had rushed through or omitted a lot of background material.  I had completely set myself up for failure.  When I went back to figure out where I had went wrong, I realized it was months earlier, when I was planning the syllabus.

Every time I teach a new class, I tell myself that I'm going to have half a dozen or so lectures ready to go by the first day of class.  Sometimes I hit this number, usually I don't.  But I've come to realize that lectures are a lot easier if you have a clear structure ready for the class.  Lately, I've become less concerned about having complete lectures prepared.  Instead of writing a few complete lectures, I try to have a fairly detailed outline of every class.  This helps me focus on the big picture.

When I have a clear direction for how the pieces fit together, the course is better.  Even if that means I don't get to the most exciting material, or have an excuse to learn something new. 

Preparing a lecture takes an unbelievable amount of time.
Truly unbelievable.  In my first semester, it took 8-10 hours to prepare each 75 minute lecture.  Add in writing assignments and exams, grading, office hours, and 2.5 hours in front of students.  Thankfully, I was only teaching one class.  But at CUNY a full load is 3 classes in a semester.  

Most of your time is spent with students who are struggling.
"Do you learn from your students?" No. I teach them.  "Are you inspired by your students?"

My PhD students are great.  They bring some exciting ideas and papers, and there's a collaborative learning that happens there.  They inspire me and drive me.  Absolutely.

Students in class, I have a very limited and lopsided relationship with.  I don't think I've ever had an office hours meeting with a student that took material from class and took it a step further.  It's almost always something to the effect of going over material from the previous lecture or exercise in more detail.  There's immense satisfaction from guiding someone to understanding challenging material, but it takes a lot of time.

Students cheat.  A lot.
I have had a student cheat in almost every course I've taught.  (This isn't unique to CUNY. Ask around.)

The cheating meeting is the most emotional human experience, I've had with anyone other than a family member or romantic partner.

The student usually cries.

The student usually tries to negotiate out of the repercussions.

Denial?  Not so much.  Most people own up to it pretty quickly.  Forceful denial is a red flag for me. Be open to the possibility that you made a mistake. Maybe one party knew about the cheating and the other didn't.  Maybe the similarities between two assignments really were random.

I've had over a dozen of these conversations.  Here's my best advice:  Be prepared. Be able to clearly explain how you know cheating happened.  Be able to point to your syllabus and university student handbook about the penalties for cheating. Absolutely document everything you can about the exchange.  Send the student an email after the fact, recapping the major points.   Expect that the student will scramble for an out -- some way to lessen the impact -- don't let them.  At this point, I have a loose script. It's almost as formulaic as a five-paragraph essay.

A. It's clear to me that you cheated on this assignment/exam.

B. Here's how I know.  (it's about here where they usually admit to it.)

C. Because of this you will be getting a zero on the assignment/failing the class/getting expelled, and I will be send a letter with this information to the department.  (Or whatever your policy is.)

Put it in writing.  Get it in writing.
Almost all I knew about teaching I learned from a video (VHS!) I was shown while at Columbia.  This was an impromptu sharing, it seemed like it was a tape that was passed around the CS department and shown by professors to their grad students.  It was a lecture by John Kender called something like "How to Teach".  It was fantastic.  It's similar to this video on iTunes.  Before you teach, watch it.

The most significant lesson I remember from this lecture was that a syllabus is a contract, an assignment is a contract, an exam is a contract.  It is your responsibility to outline the terms of this contract as clearly as you can.  This is what you will learn from this class. This is what to expect from this course, assignment, exam.  If you do this, you will get this grade.

Most of the difficulty I have had with students and disputes can be traced back to not being rock solid in the language used in a syllabus, or on an assignment, or (and this was a surprise) not putting things that were discussed in a meeting, in writing.

If you make an arrangement outside your syllabus with a student around any element of your course, shoot them an email after the meeting recapping what was discussed.  I didn't know this before teaching, but wish I had.

What now?
I've definitely become a better teacher over the last five years. I've made a lot of mistakes and I still do.

Each time I repeat a course, I toy with its structure. My boiler plate syllabus has gotten tighter.

The broader point is when I started, I was starting cold.

There must be ways, programs, seminars, etc. that aim to teach people how to teach.  I was barely exposed to any as a graduate student; I know many of my peers weren't either.  Very little was available before I was in front of students for the first time.

It's easy to gripe about a system that left me unprepared, but now i'm on the other side of the equation. I have graduate students, some of whom will go on to be professors.  What can I (and my institution) do to make sure they have the skills to be good teachers when they land tenure-track jobs (which of course they all will).

Practice. Practice. Practice.
For me, becoming a better teacher has taken practice.  I think that's the only way to learn how you are going to teach.  And there aren't enough opportunities to practice this.  (I've taught Algorithms 4 times, and Machine Learning 3.  That may sound like a lot, but it's only 2 or 3 opportunities to revise structure, lectures and graded material.)

At CUNY, graduate students teach a lot**, so many of my students will have experience in front of students.  The downside to this is 1) teaching takes a lot of time.  This means that they're not focusing on their research.  and 2) Often they're not given the responsibility/opportunity to design the class themselves.  Instead they teach a section of a larger course with a fixed set of assignments and exams.  This leaves students with plenty of experience lecturing, and leading discussions, but less experience with the mechanics of running a course (which is where your teaching lives or dies).

I think one solution might be to have graduate students prepare and teach mini-courses, complete with syllabus, and graded assignments.  These should be short, maybe 6 or fewer meetings over a month or so.  This keeps the workload more manageable compared to teaching a full course.  But it would allow students to practice structuring material, and writing homeworks and exams.  They shouldn't be offered during regular course periods, but in summer or between terms.

I think the best approach would be for this mini-course to be on the student's dissertation topic.  First of all, they'll already know a ton about it.  Second, if they go on to a tenure-track job, chances are they'll have an opportunity to reuse some of these lectures, either in a(nother) course of their own or at a conference tutorial.  Third, lecturing on a topic, and fielding questions can bring to light all the things you don't know or are unsure of.  But the practice would be useful even if it was on some other topic.

The biggest problem I see with this idea is getting the incentives right.  To teach something like this takes a lot of work, and there's little reward.  Moreover, there's little incentive for other students to take one of these mini-courses (and to do the homeworks/assignments).   MIT has a thriving IAP program with a ton of activities and minicourses ranging from the technical (some for credit) to the slightly absurd to one of my favorite things.  The IAP is well established in the MIT culture.  Can something similar be started up elsewhere?

There's no way to do this through the university registrar without a lot of bureaucracy.  However, if a department, or division, were to unofficially "bless" this kind of activity by 1) including a list of course offerings, 2) document who taught what when, and 3) conferring completion "certificates" (and 4) finding teaching space), the publicity of a program like this could encourage students to participate on both sides.

There are a lot of reasons that a program like this would fail to get off the ground, but if there were a mechanism for graduate students to get practice running courses in a relatively low-risk environment, I am confident that they would be better prepared for tenure-track positions.

I would have been.


* tenure-track
** maybe too much, but that's a different discussion

Thursday, April 24, 2014

Things I didn't know before becoming a professor (and that i'm still not very good at)


July 2009. I deposited my dissertation.

September 2009. I started a tenure-track position at CUNY.

Coming up on the close of my fifth year, I'm convinced of a simple proposition.

I was unprepared to be a professor.   

This is not a reflection on the academic preparation I had, or that my weaknesses went unnoticed by the search committee that hired me.  Neither do I think I've done a particularly bad job over the last five years.  Rather, there is a disconnect between the skills that are required to become a professor and the skills that are needed to be a good professor.

To land a tenure-track position you must:
  1. get a PhD
  2. get a solid publication record
  3. get good recommendations from important people
  4. give a good job talk
  5. be personable enough to not ruin your visit to campus
  6. get lucky (there are fewer tenure-track positions than in the past)
(If you're very lucky, you've got some funding when you walk in the door.  But this is a catch-22.  It's really hard to get funding until you're already a professor.)  

The only one of these that is non-negotiable is having a PhD.  
In order to get a PhD* you must:
  1. do good research.
  2. survive on little money and less sleep
That's it. 

Over the last five years, I've repeatedly found myself in situations where I have no idea how to do things that I am expected to do well.  I was a good candidate for a tenure-track position, but a mediocre professor. 

Here's an incomplete list of things I didn't know before becoming a professor (and that I'm still not very good** at).
  1. how to teach
  2. how to write a (successful) grant
  3. how to head up a research group
  4. that project collaboration is different from research collaboration
  5. how to manage my time
In all new jobs, there are things that you have to learn how to do, skills that get developed through practice.  But this isn't figuring out where the closest printer is, or how to fill out a TPS Report.  Most of these skills are central to the job.  There is a disconnect between the requirements to get a tenure-track job and the skills needed to do it well.  

Over the next few weeks, I'll drill down on each of these.  Hopefully, this will be some comfort to other pre-tenure faculty members, and a preview for graduate students.  It's helpful to acknowledge these challenges and the gap between what we expect from graduate students and professors.  In a perfect world, this points to opportunity for graduate programs (including mine) to provide more support to better prepare good graduate students to become good professors.  

* specific requirements vary by institution
** I've gotten better... But mostly through missteps and course corrections.

Wednesday, June 05, 2013

Deep Thoughts on ICASSP 2013

ICASSP 2013 is wrapping up today in Vancouver.  Unfortunately, I missed the last day (and sessions on speech synthesis and prosody that I would have enjoyed).  But a wedding on Saturday brought me back a day early.

I hadn't been to ICASSP before, mostly due to timing oddities and writing grants over summers rather than writing papers that would hit the deadline.  It is a very large conference.  About twice as large as Interspeech.  But the scope is also much broader.  Speech and Language work made up at most 30% of the work at the conference.  And even this is generous, including machine learning, and other work on audio.

So take this recap with a grain of salt.  I missed the last day of the conference, and my impressions are speech focused.  (I think I've described all conference recaps as blind-men-and-the-elephant problems and this one is no exception.)

Deep Learning.
OK, I pointed out that Deep Neural Nets were a "hot topic" at last years Interspeech.  It's hard to believe it's possible, but they're even hotter now.  Geoffrey Hinton gave the first plenary talk.  This was followed by an oral session called "Automatic Speech Recognition using Neural Networks", which was followed by a Special Session titled "New Types of Deep Neural Network Learning for Speech Recognition and Related Applications".  The next morning, you could attend "Acoustic Modeling with Neural Networks".  And this is just at the session level.  Even more applications of multilayer neural networks were scattered around other oral and poster sessions.  Some of these oral sessions were so crowded that people were standing along the walls and sitting in the aisles.  Nothing else that I saw received nearly so much attention.

It's easy to view "deep" learning as a silver bullet -- the next great machine learning that will solve all of our problems.  It's almost certainly not.  However, a wide array of research groups are seeing similar impressive performance gains by using deep network models for a broad spectrum of spoken language processing tasks.  This is especially true for acoustic modeling in speech recognition.  Given this, deep learning shouldn't be ignored.

Hinton's coursera course is a solid place to start. (Though resist drinking the kool-aid.  To my mind, perceptrons are bad approximations of neurons and worse approximations of the brain, and do little to advance our understanding of human intelligence.)

Another highlight
One paper which caught my attention for its simplicity came out of Google: "Language Model Verbalization for Automatic Speech Recognition".  Essentially "verbalization" is defined as a sort of inverse text-normalization. In text normalization for speech synthesis we have to translate "10" to "TEN", and "7:11" to "SEVEN ELEVEN" or "ELEVEN PAST SEVEN".  For ASR, the idea of verbalization is to convert decoding output of "SEVEN ELEVEN" into "7:11" or "7-11".  Why bother?  Well, Google (and everyone else) has big language models based on text data. You could run a text normalizer over all of this data, but the proposition here is to convert the ASR output into a form that looks more like the source material in your language model.

The Verbalizer solution to this problem is remarkably elegant.  A traditional WFST decoder can be expressed as D = C • L • G, where C comes off the acoustic model mapping context dependent to independent phones, L is the pronunciation model and G the language model.  The "Verbalized" WFST model includes a WFST V which maps ASR realizations like "SEVEN ELEVEN" to text-like realizations like "7-11" or "7:11" (and since it's a WFST it can do both simultaneously).  The new decoder looks like D = C • L • V • G.  No fuss, no muss.  Except that you have to write Verbalizer rules by hand.

The paper focused on terms involving numbers, but the framework is very extensible.  And it's great to see work coming out of Google that doesn't have Google-scale data as a prerequisite.

Meta-comment
The acceptance rate at this years ICASSP was 52%.  This means that the ICASSP and Interspeech acceptance rates are identical for the first time.  I know that Interspeech organizers have been working to lower the acceptance rate, while it sounds like there has been pressure to keep the size of ICASSP large, even at the expense of a higher acceptance rate.  IEEE (the ICASSP parent organization) is a much larger bureaucracy than ISCA (Interspeech).  There are clear expectations from IEEE about the expected revenue from hosting a conference, which translates to expectations on attendance and therefore the number of accepted papers regardless of the number of submissions.

Despite the near constant rain, I genuinely enjoyed Vancouver and ICASSP 2013.  I'm looking forward to the next.



Monday, October 22, 2012

Reading and Reconnecting

With travel finally settling down for me, but ramping up for my wife's book tour, I'm able to settle in to some long overdue reading, thinking and planning.

Also, after an exciting conversation with Dogan Can, during a trip to USC's SAIL lab, I'm trying to get more on top of sharing ideas, progress and information here.

First up: some drill-down reading from Paul Mineiro's blog post on Bagging!

Ensemble methods work too well for me to understand them so poorly, so:

  • How out-of-bag estimates can be used to get at generalization error (better than cross-validation can).  
  • The relationship between the bias-variance tradeoff and ensemble methods from this lecture.  This is a nicely framed discussion of ensemble methods that I hadn't seen before.

Lying Words: Predicting Deception From Linguistic Styles.  This paper describes a common pattern of language use in deceptive story-telling:  Less self-reference. More negative emotion words. Less cognitive complexity.  

I'm looking forward to verifying these claims on some old deception data. And taking a look at debate transcripts through this lens. 



Monday, September 17, 2012

Interspeech 2012 Recap

Portland proved to be a great venue for this year's Interspeech.  (Though people who attended ACL 2011 probably already could have guessed that.)

Setting up three simultaneous poster sessions in the parking garage may not sound like the mark of a good conference, but it was perfect.  There was loads of space between each presenter.  It allowed for all three sessions to be in the same place.  And the folks at the Hilton did a great job of making it fairly unrecognizable as a parking lot.  (In fact, Alejna Brugos didn't realize it until they were removing the carpets and "walls" on Thursday afternoon.)

Deep Neural Networks.
For "trends", there's really nothing hotter right now than Deep Neural Networks or Deep Belief Nets.  This isn't an area that I do research in, but the story goes more or less this.  Neural Networks with more than a few hidden layers don't train very well with back-propagation. Geoff Hinton and his group figured out how to overcome this limitation not too long ago.  (I think this 2006 paper explains it, but I can't be 100% sure.) Then at ASRU 2012 and ICASSP 2011 and 2012, the folks at Microsoft showed that you can use Deep Neural Networks to generate *very* useful front end features.  (Tara Sainath has a nice recap of ICASSP 2012 here.) Now, everyone wants a piece.

The field has expanded from Microsoft to include IBM and Stanford/Berkeley/Google and RWTH Aachen.  Joining them with posters on Deep Neural Nets for ASR are Tsinghua, CMU, Karlsruhe, NTT, INESC-ID, UWashington, and Georgia Tech.  At this point, there's no way to deny that this approach is receiving significant research attention.  The results seem to be holding up.  If only they didn't take so long to train...

Prominence Special Session.
I was particularly looking forward to the Special Session on Prominence.  On balance I was happy about the session.  It attracted work and discussion of prosody in a way that can sometimes feel diffuse and unfocused at a large conference like Interspeech.

I found this session to be surprising in a few ways.

It's been my understanding that "prominence" was used as a catch-all term to cover diverse prosodic phenomena including stress, emphasis, and pitch accenting.  The first surprising element of this Session came in a review of the paper I submitted to it.  The paper is on the use of automatically predicted pitch accents and intonational phrase boundaries to improve pronunciation modeling.  The review, while generally positive, found the paper to not be appropriate for a prominence session because it explored the use of "pitch accents" rather than "prominence".  I still haven't gotten a good explanation of the difference, and the reviews are blind.

A second surprise is that there seems to be a movement away from a phonological theory of prosody. Mark Hasegawa-Johnson and Jennifer Cole have been doing work over the last few years investigating how naive listeners perceive prominence.  They've consistently found that listeners respond to different qualities sometimes at different thresholds when assessing prominence.  I've found this line of research to be interesting and generally informative, but not a clear indictment of the theory that there perceptual and productive prosodic categories exist.    The panel (which I was a part of) on balance seemed comfortable with the idea that prominence is a continuous rather than categorical phenomenon.  This view was most directly expressed Denis Arnold who said approximately: focus can be categorical, stress can be categorical, while prominence is still continuous. I didn't understand this statement then, and still don't.  But again, this may be due to a different definition of prominence than I use.

The last surprise comes from finding out that there is a direction of pursuing language universals in prominence and prosody more broadly. Petra Wagner and Fabio Tamburini (the session organizers) are planning a workshop to investigate this.  In my experience, while the dimensions of prosodic variation may be used in multiple languages and some of these (e.g. increased intensity or duration) may be used to indicate prominence in all languages, it is extremely unlikely that either the communicative impacts of prosodic variation or its realization and perception are language-universal.   From that perspective, I'm not quite clear about what this line of research hopes to accomplish, but I'm curious about where it ends up.

Dynamic Decoding.
It appears that every year, I find myself sitting in on an oral session on a topic that I know very little about.  Last year it was the language identification session.  This year it was Dynamic Decoding.  I was most intrigued by this because I hadn't heard the term before.  When I asked someone what it was, they said "I don't know, Viterbi?".

I'm not quite sure this is a good enough distillation of the topic, but the papers in this session were about how to make on-the-fly (or post-training) modifications to language or pronunciation models.  This is a cool idea with clear practical importance -- how do you add words to a recognizer on a mobile device and have this appropriately incorporated into the LM and pronunciation model?  These two papers have some interesting WFST based approaches on this task.  I'll be curious to see learn more about this.  Also, if anyone has a more precise definition of this research area, I'd love to hear it.

Finally, some comments on two of the keynotes.

There were four keynotes at this year's Interspeech, two were about interesting inter/multi-disciplinary questions about how speech processing intersects with music and animal vocalization, respectively.


Chin-Hui Lee: An Information-Extraction Approach to Speech Analysis and Processing
A third was delivered by this years ISCA medalist, Chin-Hui Lee.  Prof. Lee's most famous accomplishment is MAP adaptation in acoustic modeling.  This is a researcher who spent a career treating speech recognition as a pattern matching problem.  This is a view embodied by the Fred Jelenik quote: "Every time I fire a linguist, the performance of the speech recognizer goes up".  What struck me, is that despite this view, in a talk summarizing a successful career, Prof. Lee presented a view of speech recognition that says that linguistic knowledge and speech science should be incorporated into the task.  This is an alternate perspective that has been investigated by a lot of talented researchers, including Hynek Hermansky, Jennifer Cole, Mark Hasegawa-Johnson, Alex Waibel, Hermann Ney (via speech-to-speech translation), Mari Ostendorf, Elizabeth Shriberg, Andreas Stolcke, Rene Beutler, Karen Livescu and many more (my apologies to anyone I missed).

I was struck by the evolution of perspective from someone who represents the statistical pattern matching approach to recognizing the potential importance of linguistic knowledge.


However, this talk was not so well received by some members of the audience for fairly obvious reasons.  Firstly, it over-played the importance of Prof. Lee's own contributions.  In a slide on "My contributions", virtually all major improvements to ASR over the last 20 years were mentioned including most styles of adaptation (including MAP), and virtually all major forms of discriminative training.  Secondly, it failed to recognize that the linguistic inspired approach that he was advocating for the future had been extensively researched by other talented peers.


On balance, I found it a compelling message.  In principle, it understandably rubbed some people the wrong way.


Michael Riley: Weighted Transducers in Speech and Language Processing
I should preface my comments about Michael Riley's keynote by saying that we worked together while I was interning at Google.  I'm a fan.  Michael has the rare quality of being the smartest guy in the room without letting anyone know until its genuinely useful.

The best part of this keynote was the history of the Weighted Finite State Transducer.  This was a great story that takes place largely at Bell Labs in the 90s and features Fernando Pereira, Mehryar Mohri and, naturally, Michael Riley.  This section was appropriately personal, while presenting this relevant recent history.  The WFST is so ubiquitous in speech and NLP applications that it's easy to forget that it's has a human context.

Much of the rest of the keynote felt like a 3 hour tutorial compressed into 40 minutes.  This involved showing algorithms, and example WFSTs and describing all of the things that they can be used for.  While a successful demonstration of the breadth of application, it was presented at such a pace that it was difficult to get anything out of it, if you didn't know it already.  I'd point the interested to the references found on the OpenFST page for more thorough tutorials that can be digested at your own pace.

Interspeech 2012 was successful and fun.  Portland and the Hilton (and it's solid wifi) were excellent hosts.  There was good work and as ever more than I could see.  If you have great or favorite papers that I missed, please let me know!



Wednesday, August 29, 2012

Overview of Speech and Spoken Language Processing

Here's the premise: I was invited to give a guest lecture in Advanced Natural Language Processing.   The students will get one week out of 14 focusing on speech and spoken language processing. But it's early in the semester, so there's an opportunity to give a perspective about how speech fits in to the lessons that they'll be learning in more detail later in the semester.

Here's the question: how do you spend 75 minutes to provide a useful survey of speech and spoken language processing?

My answer, in powerpoint form, can be found here.

I spent about 2/3 or so of the material on speech recognition.  I figured most students are fascinated by the idea of a machine being able to get words from speech, so let's go through the fundamentals of the technology behind it.

The remaining 1/3rd or so, I focus on the notion that speech recognition is not sufficient for speech understanding.  This a lot of other information in speech that is either 1) unavailable in text, or 2) unavailable in ASR transcripts.  The premise in this section is to convince students that speech isn't just a noisy string of unadorned words, but that there's a lot of information about structure, and intention that is available from the speech signal. What's more, we can use it in spoken language processing.

There are an outrageous amount of important concepts that get almost no attention here including but not limited to: Digital signal processing, human speech production and perception, speech synthesis, multimodal speech processing, speaker identification, language identification, building speech corpora, linguistic annotation, discourse and dialog, and conversational agents.

Would you do it differently?  I'm curious what some other takes on this problem might look like.

Thursday, August 16, 2012

AuToBI v1.3

This release to AuToBI is a more traditional milestone release than v1.2 was.  Trained models and a new .jar file will be available on the AuToBI site shortly.

There are improvements to performance that are thoroughly documented in a submission to IEEE SLT 2012.  These improvements were achieved from two sources.

First, AuToBI uses importance weighting to improve classification performance on skewed distributions.  I found this to be a more useful approach than the standard under- or over-sampling.  This is discussed in a paper that will appear at Interspeech next month.

Second, inspired by features that Taniya Mishra, Vivek Sridhar and Aliaster Conkie developed at AT&T, I included some new features which had a big payoff.  (They described these features in an upcoming Interspeech 2012 paper).  One of the most significant was to calculate the area under a normalized intensity curve.  This has a strong correlation with duration, but is more robust.  You could make an argument that it approximates "loudness" by incorporating duration and intensity.  This is a pretty poor psycholinguistic or perceptual argument so I wouldn't make it too strongly, but it could be part of the story.

Here is a recap of speaker-independent acoustic-only performance on the six ToBI classification tasks on BURNC speaker f2b.

Task Version 1.2 Version 1.3
Pitch Accent Detection 81.01% F1:83.28 84.83% F1:86.58
Intermediate Phrase Detection 75.41% F1:43.15 77.97% F1:44.43
Intonational Phrase Detection 86.91% F1:74.50 90.36% F1:76.49
Pitch Accent Classification 18.46% Average Recall:18.97 16.33% Average Recall:21.06
Phrase Accent Classification 48.34% Average Recall:47.99 47.44% Average Recall:48.31
Phrase Accent/Boundary Tone Classification 73.18% Average Recall:25.92 74.47% Average Recall:26.02

There are also a number of improvements to AuToBI from a technical side and as a piece of code.

First of all, unit test coverage has increased from ~11% to ~73% between v1.2 and v1.3.

Second, there was a bug in the PitchExtractor code causing a pretty serious under prediction of unvoiced grames.  (A big thanks to Victor Soto for finding this bug.)

Third, memory use is much lower by a more aggressive deletion of prediction attributes, and through a modification of how WavReader works.

I'd like to thank Victor Soto, Fabio Tesser, Samuel Sanchez, Jay Liang, Ian Kaplan, Erica Cooper and, as ever, Julia Hirschberg and anyone else who has been using AuToBI, for their patience and feedback.

I've been pretty lax about posting here.  I'll try to get better about it in the coming academic year.

This fall is full of travel which will lead to a lot of ideas and not enough time to work on them.