Monday, April 29, 2013

Gesture Controlled Pitch Bend

Here's a demo of a project I recently completed as part of my Cognitive Video class: Gesture Controlled Pitch Bend.


The motivation behind this project stems from my interest for the variety of sounds a guitar can produce, although I am mainly a keyboard player.
One of the expressions I have always wanted to reproduce on a keyboard are bends and vibratos. The effects are subtle, but they definitely add something to licks. Especially when a nice long bend slowly, asymptotically pass the Blue note and lands on the 5th..

There are ways to do that by using the pitch wheel featured on some keyboards. But it's a bit awkward. It's like using a whammy bar to bend a note. So I thought, why not make a keyboard a 2-dimensional device?
We use one dimension to run through all the notes - what about extracting information about the position of the player's hands along the perpendicular?

In the setup described in the video, a camera continuously tracks the position of the player's hands and sends displacement positions along one dimension to an arduino, which handles MIDI communication. The result is fairly intuitive. Slide your hand away and the bend goes up, slide towards your body and the band goes opposite.

Improvements to come:
1) I had a new idea for quick tracking of the hand
2) Make this a standalone device by using a Raspberry Pi and doing the image processing on the RPi's GPU
3) Automatic calibration to register displacements occurring over the keyboard region only.

Wednesday, February 27, 2013

Work at Tandent Vision Science


I started working at Tandent Vision Science about a month ago. I absolutely love the work, the people, and the office! The company focuses on computer vision, and recently published a release called Lightbrush. The software is amazing: in a nutshell, you can separate all the shadows in an image with one click:


I think this is incredible: so many computer vision tasks are limited by illumination variations. Lightbrush provides a fundamental first step by making the image illumination-invariant. As such, it can be used as a natural preprocessing step in all computer vision pipelines, and makes most recognition applications much simpler. Anyone with experience dealing with images for the purpose of the 3 R's (Recognition, Reconstruction, Registration) will see the incredible benefit of this software.
Besides, I am a big fan of photoshop, and I know a feature such as this one would be revolutionary! I can't count how many times I've seen awful photoshopped images due to incoherent lighting.
Apparently Lightbrush gained huge attention from texture artists and high level graphics artists (at Pixar notably) so the product is geared towards a very specialized crowd. Too bad, I personally think the average user would have a blast playing with this product.
I've been working on improving the machinery - can't talk about it! - and it still blows my mind that this is possible with minimal user input... There are many aspects that I have to take into account in my tasks, such as keeping the user in mind with respect to user interaction complexity, speed (texture artists work with gigantic images).

I'm working with 2 very friendly Carnegie Mellon alumni. We each have our separate offices (with a REAL DOOR - something I missed at other workplaces I spent time at) and has a very cozy homey feel, much different from the old functional cubicle. I'm definitely loving the work here.


Automated caller

It's been a while since I updated anything here. My schedule for this new semester is fairly busy: 4 classes, research and an internship barely leave enough time to cook.
The script below is something I wrote back in December when contact information from a certain religious hate group was hacked then published online. It successively and continuously calls all the numbers on a text file (in the file below, the numbers are just stored in an array). I used mechanize to fill forms on website such as findmyphone.com and dial the numbers.


I was debugging the script using my number and had to leave in the middle of it running... I came back to 57 missed calls and voice messages. 

Saturday, December 15, 2012

HMMs

Here' s a really good explanation, through concrete examples, of Hidden Markov Models.
[Credits: Professor Moore, Carnegie Mellon SCS].


Saturday, December 8, 2012

New university webpage

Crunch mode over. I just created my personal webpage on Carnegie Mellon's ECE servers.
Links to a project paper on automating Horizontal Gaze Nystagmus (part of the field sobriety tests performed by law enforcement in the US) and the final paper on using Twitter to predict users' political affiliations are on there.

My girlfriend thinks I really should learn HTML5 - That might be my Christmas break project.

Sunday, November 18, 2012

Twitter API & Ruby

     Last Monday was the day the mid-semester report was due for our Machine Learning project. That means we went into full crunch mode the week before. And that means we went into a whole bunch of changes from our original idea.
First change, the new project is not so heavily computer vision oriented: we want to classify Twitter users on their political affiliation. This has direct relevance in the context of the 2012 presidential elections. Can we predict a particular user's vote?

Dataset
     We collected Twitter user IDs through the Twitter API in Ruby. One of the major issues we had to work around is to get a ground truth - which unfortunately, is not provided in Twitter profiles.
To address this labeling issue, we created an approach using some domain knowledge, which ensures that we have a label for each of our user by inferring their political affiliation. We queried users who follow President Obama; while it might reflect some sensitivity to Democratic convictions, we also required that a user also follow several of the following list: Joe Biden, Stephen Colbert, Jon Stewart,  Bill Clinton, Hillary Clinton, Al Gore...
Conversely, a 'Romney' label is applied to users that follow Paul Ryan, Sarah Palin, the NRA, Bill O'Reilly, Rush Limbaugh, Glenn Beck...

Algorithm
     Once the user IDs were collected and labeled, we fetched their tweets, age, location, relationships (followers/followees)...
The idea is to apply a bag of words approach to each user - hence, the resulting histogram are part of the feature vector of each user.
What about the bins? We created a long list of "strong" words (in regex format) - such as "Gun Control", "Obamacare", "Occupy Movement". We believe these words polarize the tweets, thus capturing valuable information to cluster users.
The next step is to run Lloyd's algorithm on the instances in the dataset (aka k-means). Because we know that it is heavily dependent on the initialization, we have a method to have more coherent and more intuitive initializations. Each value of k is associated with an objective function we are trying to minimize, which reflects the intra-cluster variance. We swipe over a range of k, storing the objective function for each, which allows us to plot Objective function vs k.
This plot provides us with some insight as to which optimal value for k we should use. When each instance has a feature vector of length 150+, it is impossible to visualize the data and determine the number of clusters visually.
One benefit of using k-means is that we don't need to carry the dataset for classification. Once the cluster centers have been determined, that is all we need to classify a new instance (unlabeled user).
The last step is to perform dimensionality reduction to express the data with only the most relevant features. PCA, LSI in topic models are paths we will explore at that point.

Some issues we've had to deal with is getting around the restriction on the number of queries set by Twitter. A single authorization token (on Twitter for developers) will provide a maximum of 500 queries/hour. To get passed this, we created a bunch of authorization tokens, which we cycle through.

Here's what a single token authentication looks like, followed by a user timeline query:
require 'twitter'
require 'json'
require 'pp'
require_relative 'Token_Nico.rb'

YOUR_CONSUMER_KEY= "mnnH8LoWM7#########"
YOUR_CONSUMER_SECRET ="FTYr8xgRdTyMEACPEO9Jfxl##################"
YOUR_OAUTH_TOKEN = "23894652-NeIDQ4JeHMofJxldF#######################"
YOUR_OAUTH_TOKEN_SECRET= "UgvOuWpaTTnhKKLpiHz9##################"

@client = Twitter::Client.new(
                              :consumer_key => YOUR_CONSUMER_KEY,
                              :consumer_secret => YOUR_CONSUMER_SECRET,
                              :oauth_token => YOUR_OAUTH_TOKEN,
                              :oauth_token_secret => YOUR_OAUTH_TOKEN_SECRET
                              )
# get timeline from id
tw_data= @client.user_timeline( id.to_i, :count=>200, :exclude_replies=>false, :include_entities=>true)

'Token_Nico.rb' contains a class definition as follows:

class ClassToken
    def self.re_configure(i)
    pp "reconfigure i=#{i}"
        pp Token[i][:consumer_key]
        @client = Twitter::Client.new(
                                      :consumer_key => Token[i][:consumer_key],
                                      :consumer_secret => Token[i][:consumer_secret],
                                      :oauth_token =>Token[i][:oauth_token],
                                      :oauth_token_secret => Token[i][:oauth_token_secret]
                                     
                                      )
        return @client
    end

    def self.howmany?
        return Token.count
    end
end
 Token is an array of token arrays. The idea is to loop through the authentication keys until either we get the user timeline, or we determine that the timeline is protected (profile set to private).

Sunday, October 21, 2012

Collecting data: Twitpic Scraper with Matlab


For my Machine Learning course this semester, the project I will be working on is Context-Based Object Recognition using Twitter. We would like to use the associated information about twitted images (author gender, hashtags, caption, comments etc) to try to improve recognition. Those extra features will be used as a prior, providing a context for the classification pipeline.

First step is to create our dataset. I wrote a Matlab script that scrapes the desired features from twitpic with a query for the keyword "pet". Here's a first result:
Original twitstream:

and collected data in Matlab (title of each subplot is the associated caption, which can be too long and overlap):
 The data is stored in a structure that currently has 3 fields: {image, caption, hashtag}. Here's a more detailed view where we can see the extracted #hashtag:

 Here's a link to the script I wrote: Twitpic Matlab scraper. Note that it uses the Twitpic API to get an XML response. We're looking at collecting a dataset of around 1000 instances.

One thing I have yet to address: it seems a lot of tweets are in non US-ASCII character set (for a "pet" query, a lot of the captions were in Japanese). So I will need to modify the script slightly to treat those tweets differently.


Edit: Older version of the script would crash if the visited profile was recently created (missing information such as post history) or if there was a video instead of an image.

Here is a more robust version of the Twitpic Matlab scraper. The user can now specify how many pages to scrape, and which tag to look for.