Go to: LING 1330/2330 home page  

Exercise 6: Boy or Girl? Movie Good or Bad?

The goal of this exercise is to learn how to build a Naive Bayes classifier. The NLTK has two examples: Name gender classifier which we covered in class (save shell session here), and Movie review classifier, which is the focus of this exercise.

Movie review sentiment analysis

Try out document classification on movie reviews by following this NLTK Book section, operating in Python shell. Since we are dealing with positive/negative opinions on movies, this task is a form of sentiment analysis. Details:
  1. First, review the "Name Gender Classifier" we covered in class, if you haven't already. (You don't have to submit this part.)
  2. Start out by exploring the movie reviews corpus and familiarizing yourself, which the book didn't do. Find out how big the corpus is, how many reviews there are, and how many of them are positive/negative. Take a look at a positive (or negative) review to get a concrete sense of the content.
  3. You will notice the code in the book is pretty dense with lots of list comprehension. If you find a code block confusing, focus instead on the end result: what the newly built data object looks like, and how it's structured.
  4. Because of random shuffling, your "most informative features" list might not look exactly like what's shown in the book. So, don't be alarmed if Mr. Matt Damon is missing from your list. Don't stop at top 5 features: try 20 or more.
  5. If you're done with what's in the book, it's time to try something new. See how the classifier classifies this short and fake movie review.
     
    >>> myreview = """Mr. Matt Damon was outstanding, fantastic, excellent, wonderfully 
    subtle, superb, terrific, and memorable in his portrayal of Mulan."""   
    >>> myreview_toks = nltk.word_tokenize(myreview.lower())  # lowercase, and then tokenize
    >>> myreview_toks
    ['mr.', 'matt', 'damon', 'was', 'outstanding', ',', 'fantastic', ',', 'excellent', ',', 
    'wonderfully', 'subtle', ',', 'superb', ',', 'terrific', ',', 'and', 'memorable', 'in', 
    'his', 'portrayal', 'of', 'mulan', '.']
    >>> myreview_feats = document_features(myreview_toks)     # generate word feature dictionary
    >>> classifier.classify(myreview_feats)    # classify
                  ??              
    >>> classifier.prob_classify(myreview_feats).prob('pos')  # probability of 'pos' label
                  ??              
    >>> classifier.prob_classify(myreview_feats).prob('neg')  # probability of 'neg' label
                  ??              
    >>> 
    
  6. This time, change "Matt Damon" to "Steven Seagal" (IMDB profile) and see what happens.


SUBMIT:
  • A saved shell session as a .txt file, edited to clean up messy bits and to include your notes/comments