Monday, January 24, 2011

Cleaning Azure Marketplace Data with PowerPivot and Predixion Insight

image
Thanks to some connections at Microsoft I was made aware of some very interesting data that was published on the Windows Azure Data Marketplace by Practice Fusion – deidentified medical records of 5000 patients.  Of course, I had to take a look.  What I found was a very compelling and interesting set of data, that was more interesting when looked through the lens of Predixion Insight.  Of course, as with any data that you may get from anywhere, there are first some issues in the data itself that need to be “corrected” before used for any analysis.
In this post I’ll show you how to get that data, explore a little bit of it, and clean up one particular part.  There’s lots more to do, but this is just getting your feet wet.
In order to follow along at home, you will need Excel 2010, PowerPivot (get here), Predixion Insight for Excel (get here).

Step 1 – Loading the data

After you have all that software installed and running on your machine, you will want to go to the Practice Fusion Data page on Windows Azure and subscribe to their feed.  If you haven’t subscribed to any the data market before, you will first have to agree to the MS terms of service to subscribe to the data market.  After that you will have to agree to Practice Fusion’s terms of service and that will give you access to the data.  Ignore the button there that allows you to load data directly into PowerPivot.
The most efficient way to get all of the data is to go directly into PowerPivot, so launch Excel (or just create a new workbook since you probably already have Excel running anyway) and click on the PowerPivot button to get to the PowerPivot window.
Get External Data
PowerPivot Ribbon Section
Inside the PowerPivot window you will need to click the “From Azure Marketplace” button in the Get External Data chunk of the ribbon shown below.  If you’re saying “Hey! I don’t have that button!” don’t be afraid.  Microsoft quietly updated PowerPivot at least twice since it was released – go and download the latest release and try again – trust me – it will be there.
The From Azure Marketplace Button launches the import wizard where you want to paste in the root URL for the data feed which is: https://api.datamarket.azure.com/Data.ashx/PracticeFusion/MedicalResearchData/
Also you will need to enter your account key, which, up to this point, you probably didn’t know you have.  However clicking on the Find button will launch a web page displaying your unreadable key that you can copy and paste here.
Specifying Data Feed
Table Import Wizard – Specifying the Data Feed
After that, select all the tables and click finish.  Loading the data from the marketplace takes a little longer than you would expect – not “come back tomorrow” kind of long, but definitely the “I can’t be late for my daughter’s junior high school recital and if I take the laptop, I’ll lose my wireless connection” kind of long.  So click finish and go to lunch, or get a coffee, or catch up on reading the Predixion Insight documentation or something.
Selecting Tables
Table Import Wizard – Selecting all the Practice Fusion Tables

Step 2 – Initial Data Exploration

One of the somewhat interesting things about loading data from an unknown source is that you have absolutely no idea what you’re looking at.  This is where you can leverage the awesomeness of Predixion Insight’s Profile Data tool.  This tool provides a quick summary of the data in an Excel worksheet or PowerPivot table so you can at a very minimum determine if the data makes any sense at all.
imageProfile Data Command in Predixion Insight 
To run Profile Data, switch back to the Excel interface, click on the Insight Analytics ribbon and choose Profile Data under the Explore Data menu button.  In the Profile Data Wizard, you just need to select that you are choosing PowerPivot and the source of the data and which table you want – in this case, let’s use the SyncChart table which documents the basic patient measurements over a series of doctor’s visits.
Profile Data Select DataProfile Data Select Data Page
Next you choose the type of statistics – I really like using “advanced” statistics because it makes me feel smarter ;), but for initial exploration the default “basic” stats will suffice, so stick with the default and just click “Finish”
Profile Data Options
Profile Data Options
As a result, Predixion Insight will scan the table and produce a report on all of the columns in the PowerPivot table, as well as a handy index sheet to easily find the report in your workbook (which at this point should contain the report, the index, and three blank sheets, so it’s no biggie, but you know it will get more populated as you go on…).  I can’t recommend strongly enough that, for your own sanity, you go ahead and run Profile Data on all of the tables in PowerPivot (see, that index will come in handy very quickly….)
In any case, let’s have a look at what Profile Data discovered.
image
Results of Profile Data on the SyncChart table
First of all, Profile data did us a great service by simply telling us how many rows and columns are in the table – since we just randomly (well, not quite) pulled some data from the internet we had no idea what was in there, Profile Data let us know that we have 59,193 rows and 12 columns, and that the column ChartGuid is likely a key since it uniquely identifies each row (i.e, there are 59,193 unique values of ChartGuid).
Next the Profile Data report breaks the results down into a Continuous section and a Discrete section.  The Continuous section contains information on columns that contain numerical data.  The Discrete section contains information on columns that contain non-numerical data, or numerical data that “look like” they could be considered categorical, such as a column that contains just 0’s and 1’s.
One thing that stands out fairly immediately is that there is a lot of missing data in this table.  All the continuous columns, save VisitYear, have over 23,000 blanks, and the HeartRate column is completely empty – i.e. it has 59,193 blanks in 59,193 rows.
Feel free to examine the resulting report for interesting pieces of information but for expediency’s sake, I want to point out some data that just looks wrong.  Take a look at the ranges of values for Height, Weight, and BMI
image
Focus on Height, Weight, BMI
Now, not having any background in the data, we don’t know what units height and weight are measured in, but regardless of the units, the ranges are simply too extreme – given inches, feet, or centimeters, you simply don’t have the range of heights from 5 to 120 in human existence!  Similarly with weight!  BMI is unitless, but still with the range from 0.07 to 5511 you have quite the unnatural Jack Sprat situation going on here.
Looking at the results for the height column a little further indicates that there is a mean of 65 and standard deviation of about 5.  This would indicate (to me, at least) that the intended units are inches, since the average height of humans is around 65 inches and not 65 centimeters or 65 feet or miles or whatever.  So I’ll go ahead and decide that it’s inches.  It would make sense that the weight would be in pounds since the height measurement is not using the metric system either, but looking at the average weight of 180 indicates that the value, while still high (this is America, anyway) isn’t likely to be kilograms or stone or other values.
Still, we need to figure out what is going on with those other values, so let’s move on to Step 3.

Step 3 – Cleaning up the data

Cleaning up data like this is never automatic – you need to use tools and apply your judgment.  Just like we were able to use the Profile Data tool to help us understand the nature and deduce the units of the data, we need to apply the same common sense treatment to cleaning the data.
The first question to ask is what do you really want to do with the erroneous data?  Since these are medical records recording actual values of patients, the safest thing to do may be simply to null out the data for values that are out of range.  However, since there are several rows for each patient, and people (once fully grown) generally don’t change in height too much, it may be OK to simply take a height measurement from another chart reading for the same patient.  It really depends on how you’re going to use the data, so I’ll perform both methods.
Let’s take the “whack the bad data” approach first.  To do this, I’m going to use the Outliers tool from Predixion Insight which is under the Clean Data menu button on the Insight Analytics ribbon.
image
Clean Data/Outliers
Selecting this tool brings up the Remove Outliers ribbon which allows you to truncate the extreme ranges of numerical values or remove categorical values that occur infrequently (or even too frequently) from your data.  First you need to select that you are interested in the PowerPivot table called SyncChart (as above) and then you select the column from which you want to remove the outliers.  You can select via the drop down or by clicking on the column preview.
image
Outliers Column Selection
Next this brings up a chart of the values in the Height column.  The higher the curve the more values are in that range.  Personally, I like to increase the resolution of the chart to the maximum (100) so I can get the best view of the data – this sometimes doesn’t work well if there are a small number of data points, but this data set is sufficient.  As you can see most of the data lies in the middle of the range with extreme values tapering off quickly.  You can mouse over any point on the cart to see the range that point represents and how many rows have values in that range.
image
Remove Outliers Chart
To remove the outliers you simply drag the thumbs on the slider above the chart to whatever cutoff point you wish, or type in the limits into the text boxes directly.  When you click next you get to the “what do I do with this?” screen.  Again, the answer really depends on your application.  For some problem spaces, it may be OK to set the extreme values to the limits or to the mean.  In this case, we’re simply going to null out the value – the default option – and click next.
Remove Outliers Options
What do we do with a drunken data entry operator?
Which leads us to the final screen  - name the result column – it’s good to pick a descriptive name.  Finishing the wizard creates a calculated column in PowerPivot that nulls out any data that’s out of range.  This has an added benefit in that if you refresh the table with new data, any new out of range data will automatically be set to null.
Select Destination
Finishing the Remove Outlier Wizard
Inspecting PowerPivot, we see that Predixion Insight injected the following DAX expression:
     =IF(AND('SyncChart'[Height]>= 48.99, 'SyncChart'[Height]
      <85.28),'SyncChart'[Height],BLANK())

This expression simply copies the in-range values and replaces the out of range values with blanks.
Before looking at the results, I want to try the other method of removing outliers I proposed – using the average of a patient’s other visits.  To do this I used the following custom DAX expression:
   CALCULATE(AVERAGEX(SyncChart, SyncChart[Height]),
       ALLEXCEPT(SyncChart, SyncChart[PatientGuid]),
       SyncChart[Height]>48.99,
       SyncChart[Height]<85.28)

This expression basically says, calculate the average of the Height column by first removing all filters except for a filter on the current rows PatientGuid Column, and also exclude any rows where the Height is out of range.  I copied the IF expression created by Predixion Insight and replaced the part that says “BLANK()” with this custom expression.  I renamed the new column “Height with Outliers Replaced”
The next thing I want to do before looking further is calculating the BMI.  The BMI was included with the data, but is really a derived column.  The formula for BMI using inches and pounds is BMI=Weight * 703/Height^2.  In order to protect against divide by zero errors, I put an IF in the DAX expression to check for blanks and created two columns like this:
     =IF(ISBLANK(SyncChart[Height with Outliers Removed]),
            BLANK(),
           SyncChart[Weight]*703/POWER(SyncChart[Height with Outliers Removed],2))

(Obviously one had the Outliers Replaced version of the column).
To see the impact of the results, I re-ran Profile Data against the SyncChart table with the new columns. Here are the results:
Profile Data Results
The first thing I noticed is that the Height with Outliers Replaced column has no blanks – this means two things – one, that every patient has at least one doctor’s visit where they have a height measurement, and two, I should have checked for ISBLANK in my expression to calculate that Height value!
The other things to notice is that the fixed height values are much more in line with reality  - we’ve eliminated the pixies and the hill giants (frost giants are much taller).  Also, the recomputed BMI values are starting to be more realistic as well.  In fact we reduced the maximum BMI from over 5500 to 250 by only removing 317 values from 59193 rows – about 1/2%! 
To complete the exercise and ensure that the BMI values were reasonable we would have to clean the Weight column as well, which we do by running through the Remove Outliers wizard again.  However, there’s actually one more trick.  When looking at the outliers for the Weight column, we get a chart that looks like this:
Specify Thresholds
This occurs when you have a range of outliers that includes extremely large outliers that obscures the resolution of the data.  To get a better view of the data, you can click the logarithm button – circled below – and get a much more usable visual – I also increased the resolution:
Specify Thresholds Log
I removed the extreme outliers (rather arbitrarily choosing a range of about 33 – 510 lbs, which is still probably pretty extreme) replacing them with nulls.  I didn’t try to do the “replace with other visits” trick, since people’s weight can fluctuate quote a bit from visit to visit.
Creating a new BMI column and re-running the Profile Data gave me results like this (other rows hidden):
Profile Data
You can see the range of BMI values went from an original 5511 to a more reasonable 78.  It’s very likely (pretty much guaranteed) that I could have more tightly truncated both the height and weight variables to make them more inline with reality, but this is a good start. 
Some deeper inspection of the data may be necessary to find the correct boundaries, and of course it always depends on how you are going to use the data.  For predictive purposes, it generally is better to eliminate extreme values even if they are valid since they can skew the results.  For traditional BI reporting – e.g. “how many visits were by people over 9 feet tall” – you generally want to keep any valid value, no matter how extreme.  In either case, however, you do want to eliminate invalid values.  Luckily with this data set, by combining common sense with the tools at hand we’re able to carve away at the bad data to make the remaining data useful for analysis.
That just may happen in a future post… ;)

Thursday, December 9, 2010

Predixion Insight Update

image

 

We just launched a new, incremental version of Predixion Insight with a whole bucket load of new features – some subtle, some more obvious and in your face.  Actually – many you shouldn’t even notice – just have a feeling of the product being “better” all around.

In this posting I’ll just give a quick overview of some of the new additions and later I’ll get around to providing details on all the goodies.  It’s better if you just try them our for yourself anyway – if you are a current user, just log in and you’ll be prompted to update, and if not what are you waiting for?  (One great new “feature” is that we’ve extended the free trial to 30-days, to give you ample time to check everything out!)

So I’m going to tell you about new features in my trademark stream of consciousness – no particular order kind of way.

Normalization

image

The Normalization allows you to normalize data in Excel or PowerPivot by Z-Score, Min-Max, or Logarithm.  It’s actually really cool because you can normalize all the data at the same time instead of column per column.    Also, if you’re normalizing PowerPivot data, you can conditionally normalize so you can get Z-scores by group.  Way cool.

 

 

 

Explore Data

image

We added additional data exploration options so the “Explore Data” button became the “Explore Data” menu and the previous “Explore Data Wizard” became the “Explore Column Wizard”.

Profile Data provides descriptive statistics about your Excel or PowerPivot data.  I won’t say too much about it here since Bogdan already wrote up an excellent post just on this feature! 

Also, Explore Column has an added button to allow you to look at your data (and bin it!) in log space.  All your data scrunched up on the side?  Click the “log” button and see it spread out with a nicer, friendlier distribution.  This button also shows up in the Clean Data/Outliers Wizard as well.

Getting Started

image

I guess I really should have put this one first, despite this being stream of consciousness.  Anyway, noting that customers were having a hard time finding our sample data, help and our support forums, among other things, we added a handy (and pretty) Getting Started page.

 

PMML Support

image

In the Manage My Stuff dialog you can import models from SAS, SPSS, R, and any other PMML source.  You can then use all of the model validation tools and the query wizard to score and validate against data in Excel and PowerPivot.  Way to leverage your extend investment in predictive analytics tools to the Excel BI user base!

Insight Index and Insight Log

image The Insight Index is an automatically created guide to all of the predictive insights and results you create with Predixion Insight.  This guide provides descriptions of every generated report along with links to each report worksheet that are automatically updated when you rename worksheets and change report titles.  The Insight Log automatically maintains a trace of all Predixion Insight operations that generate Visual Macros.  These operations can be re-executed individually or copied to additional worksheets to easily create custom predictive workflows.  Both of these features can be turned on or off in the options dialog.

VBA Connexion

Predixion Insight now supports VBA programming interfaces allowing you to create custom predictive applications inside Excel.  It’s like super hard – for example, look at this excerpt I wrote for a demo that applies new data to an existing time series model so you can forecast off of a short series:

image

Oh, wait – it’s not hard – it’s easy!

Visual Macros

image

While we’re on the topic of Visual Macros, I should mention that Visual Macros now support all of the Insight Now tasks as well as the Insight Analytics tasks.  This means that you can easily encapsulate any operation we perform on the server in a macro and use it to string together your own workflows.

Other stuff working better

There are always little things here and there that are better as well that you don’t notice until you get there – the Query Wizard is cleaner and easier to use.  Exceptions – even ones caused by data entry – give you an immediate way to provide feedback directly to the dev team.  And I can’t even begin to talk about how much better our website is!  (Mostly because I’m out of time to write this post – HA!)

Really – try it out and let me know what you think.  If you’ve tried it before go check out what’s new – if not, now is a wonderful time to do so, it’s lookin’ pretty good.

Tuesday, September 14, 2010

And launched! (plus some secrets)

Yesterday we officially launched our first  offering from  Predixion Software – titled “Predixion Insight”.  For those who haven’t been following along, Predixion Insight is a cloud predictive analytics service access through an embedded Exceyippeel client (Predixion Insight for Excel) that has absolutely no infrastructure or procurement friction and works with Excel 2007 and botht he 32-bit and 64-bit versions on Microsoft Excel.  Oh – and it’s also directly integrated with Microsoft’s new PowerPivot offering allowing powerful analytics, business modeling and now, predictive analytics, right in the Excel working environment.

We actually closed down the beta and turned on Insight’s lights a week ago today.  Coming from a “packaged software” background it was interesting that shipping software in a cloud environment all boils down to switching a DNS entry on GoDaddy.com and we’re up and running.  We actually had one customer forego the free trial and become our first paying customer on day one!  A lot of our beta customers came back to take advantage of the free trial (sign up here if you haven’t already) and our service has been humming along making predictive magic for everyone quite nicely.  I do thank all of our beta customers for helping out and finding issues based on configurations and connection performance that would have been difficult for us to find out on our own – we also managed to incorporate quite a bit of feedback from the beta into the product.

On Friday we celebrated the launch in the dev offices with a toast of Pyrat Rum and then heads down to keep the Predixion train going.  Personally I’ve been busy with Simon, our CEO, demonstrating the product to industry analysts – 20 so far, and many more to come – and getting great feedback and some exciting responses (first published notes here, here, and here), while designing elements of our Enterprise offering as well as some exciting incremental functionality.

One interesting thing about any product release is always what is the last feature to make the cut.  What is that last piece of functionality that you just need to put in or that you want so badly that it gets in no matter what.  In our case it’s all in the task pane of Predixion Insight for Excel – the task pane was definitely the runaway surprise success story for the beta – one participant even responded “I wish all software worked that way.”  Given how happy people were with this feature, and we always wanted to add a little more functionality, we made sure that search and filter made it into v1.

image With this feature you can type arbitrary text and Predixion Insight for Excel will search all fields (including extended info) of each task and only display the results, or you can filter based on task type or items that are expiring soon.  Neat, huh?  Ok, maybe just “ok”.  Anyway, since you’ve already downloaded, installed the software, and signed up for a free trial, you may have noticed that if you select one of those filtering options it actually places something in the search box like this:imageThis is totally meant to imply that there can be other things you could “type” into the search box to filter your tasks, and as of today, this is the only place you’ll ever find out about what they are.  Search tags usage is simply “tag:value”  if the engine finds the “tag:” starting the search string it uses it – plain and simple – it’s not even case-sensitive.  The super-secret Predixion Insight for Excel search tags are:

Tag Value Description Example
tasktype InsightNow
Modeling
Test
Query
Management
Returns tasks of the specified task tasktype:Query
expires integer hours Returns tasks expiring within the specified number of hours expires:48
created integer hours Returns tasks created within the specified number of hours created:2
duration integer minutes Returns tasks with a duration less than the specified minutes duration:1
completed (optional)
true
false
Returns tasks that have completed, or not completed:
results (optional)
true
false
Returns tasks that have results to download, or not results:
status succeeded
failed
pending
Returns tasks with the specified status status:pending
tag any string Searched only the “tag” field of the task – i.e. the large bolded text tag:Profit Chart

You can add a bang (!) to the beginning of the tag to return any results not specified by the filter – for example you can get all the tasks created in the last hour by using the string “created:1”, and you can return all the tasks not created in the last hour by using the string “!created:1” – very handy indeed.

Oh, and just for a bonus, you can edit the “tag” (large bolded text) of any task for  your own personal edification – that is, you can change “Profit Chart” into “My Superbad Market Mayhem Chart” and then find it with the search “tag:Superbad”

Anyway, that’s what is there for now – who knows what may come as Predixion Insight grows…..

Wednesday, August 18, 2010

Closer to liftoff…

This week we officially launched our public beta, and itimage was one of those moments where you’ve pushed so hard running on adrenaline that when you’ve reached that summit you collapse because you can finally sleep a good sleep – if only for a moment.  With the beta launch we have people from around the world enjoying predictive analytics in the cloud via Predixion Insight.  It’s exciting watching from behind the scenes as customers launch asynchronous predictive tasks ranging in size from a few kilobytes to 100 megs.  The machinations of a Rube Goldberg contraption comes to mind as the pieces of the system coordinate – a user presses a button causing their data to be launched to the image cloud while simultaneously they are automatically provisioned across an array of servers.  The data shuttled seamlessly and invisibly between tasks on their behalf being dissected and analyzed before being dropped into a predictive report right back on their desktop. 

The movement from development to beta deployment is really wonderful for me personally.  If you haven’t yet seen it, go to our website and click on play video to get an overview of the company.  If there’s one word I can say about that piece of “marketing,” it is that it is sincere.   Go ahead – go watch it.  This is something I’ve been working toward a long time.  We’re releasing a version 1 product and we have a lot to do to fully reach our goals, but right now, any user, anywhere, can access powerfulimage easy-to-use predictive analytics without having to jump through hoops for procurement, acquisition, installation, management, etc. etc. etc.   By creating an Excel-native, subscription-based predictive service, we’re taking the traditional barriers barring people from even opening the door to predictive analytics and slashing them to the ground.  

So I was going to write a longer post explaining some more details about the product, but you should try it now (and anyway, Bogdan already wrote a great post with some feature details)  You can watch a demo that gives a lightning fast overview of the product here and then go download and enroll for the “free trial” beta.  I have it on good authority that there may be some interesting beta events for accomplished users, and you still have the opportunity to get in early, so don’t wait!

Sunday, August 1, 2010

Predixion on the brink….

   I officially started my career as “founding CTO” with Predixion on January 6, 2010, and now, just 7 months later we are on the brink of launching the VIP beta of our new product and service, Predixion Insight, on August 2.  WithLG1 a development team of only 5 people we’ve created, what I think, is a truly disruptive entry in the predictive analytics space, and we’re just getting started.

It’s been a very exciting time – meeting the co-founders of Predixion, deciding to venture off from Microsoft to start something new based on the ideas developed over the past several years, recruiting the best development team you could ask for, filming corporate videos at my house, meeting with customers, partners, and venture capitalists – there hasn’t been a boring day yet!

This last week we’ve moved to a new office space in Redmond and wrapped up the bits for our VIP beta.  This beta is limited to only 12 select people.  We ran two incredible online demos and some feedback we received: “let me say that I loved it”, “Can't wait to play with the product!”, “Based on what I saw yesterday, Insight is more like a coral reef than a warm bath!”

Anyway don’t be worried that you will be left out because you’re not part of our VIP beta – we are quickly filling up our next phase of the beta offered on a first come-first serve basis.   This phase will be launched on August 16th and you can sign up on our website.  We’ve been working hard and fast at making a product that you can use immediately, every day, without boundaries, and we’re on the brink of delivering it to you.  Over the next few weeks we will be creating collateral materials that make using Predixion Insight even easier.  Stay tuned, true believers, you’ll like what we have coming!

Friday, May 7, 2010

Bootstrapping Windows on GoGrid – getting your admin password on the box.

 I spent a lot of time this week working on trying to get our service running on GoGrid as a potential alternative to Amazon’s EC2.  They jury’s still out, but they seem to offer better hardware for the price.  There are a lot of other pro’s and con’s between the two services, but maybe that’s a subject for a future article – maybe after we make a final decision!  The nature of our service requires that we can perform on-demand machine requisitioning and provisioning.  Using Amazon’s EC2, certain aspects were easier than GoGrid, due to the nature of the way they handle server images “AMI” in Amazon lingo, “MyGSI” in GoGrid.  In short, the nature of the sysprep step performed on a newly provisioned machine at GoGrid causes some problems with certain services we need to run and user accounts we need to provision.
Part of the issue has to do with the way GoGrid provisions administrator passwords – on a newly provisioned machine, the administrator account will have a new password, which you would expect, but also, any additional administrators you create are on the image aren’t valid after provisioning.  So the GoGrid-provisioned password is pretty much all you have.  This is OK if you can interactively logon to the machine after provisioning, but not so OK if you want to do this automatically.  To solve this problem, I came up with a method to fetch the admin password from GoGrid itself from the machine after launch.  We trigger this via a web service call after the machine is launched, but presumably you could do this on a startup event as well – I haven’t experimented with that as of yet, but presumably it should work.
The difficulty in the solution is simply due to the limited information you have about your machine from your machine.  The basic approach is to call the GoGrid API to get the list of passwords from all your machines, and then find the password that matches the public IP of your machine.  In order to use this code, the first thing you need to do is to go to your GoGrid account page and add an API key which you will use to securely interact with the GoGrid service.  The type of API key should be System User, as that is required to fetch passwords.  This key will be embedded in your code on the GoGrid image, so you should take necessary steps to protect it.
In this solution I use the GoGridClient class from the GoGrid Wiki Documentation – copy that code and specify your api_key and shared secret.
The first task is to write a function to get the passwords from GoGrid (we wrote the GoGridIPType and GoGridIPState enums – they contain the values in the code):

public static string GetPasswordsRaw() // returns the raw XML as provided by GoGrid
{
    string returnValue = String.Empty; 
    try
    {
        GoGridClient grid = new GoGridClient(); 
        System.Collections.Hashtable parameters = new System.Collections.Hashtable();
        parameters.Add("format", "xml");
        string requestUrl = grid.getAPIRequestURL("/support/password/list", parameters);
        returnValue = grid.sendAPIRequest(requestUrl);
    }
    catch(Exception)
    {
    } 
   return returnValue;
}

After you have this function, you need a function to get the list of ip addresses from your machine and compare it to the ip addresses from GoGrid.  The function first grabs all of the ipaddresses from the local machine and then uses Xpath queries to isolate and iterate the password objects from the GoGrid response.  Then it uses more Xpath queries to grab the ipaddress and password from each object.  Finally it checks to see if the ipaddress matches any ipaddress on the machine and returns the associated password.

private string GetAdminPassword()
{
    // Fetch ip addresses for the local machine and store into a list
    List<string> ipaddresses = new List<string>();
    System.Net.IPHostEntry IPHost = System.Net.Dns.GetHostEntry(System.Net.Dns.GetHostName());
    foreach (System.Net.IPAddress ip in IPHost.AddressList)
    {
        // Only take the IPv4 addresses
        if (ip.AddressFamily == System.Net.Sockets.AddressFamily.InterNetwork)
        {
            Report("Found ip: {0}", ip.ToString());
            ipaddresses.Add(ip.ToString());
        }
    } 

    // Get the password information from GoGrid and load into an XML document
    string xml = GetPasswordsRaw();
    XmlDocument d = new XmlDocument();
    d.LoadXml(xml); 

    // Use Xpath to select the "password" objects
    string path = "/gogrid/response/list/object[@name='password']";
    XmlNodeList nodes = d.SelectNodes(path);
    foreach (XmlNode node in nodes)
    {
        // Extract the password and ipaddress from the password object
        XmlNode pwdnode = node.SelectSingleNode("attribute[@name='password']");
        XmlNode ipnode = node.SelectSingleNode
            ("attribute[@name='server']/object[@name='server']/attribute[@name='ip']" +
             "/object[@name='ip']/attribute[@name='ip']"); 

        // API Key passwords will not have an ipnode
        if (pwdnode == null || ipnode == null)
            continue; 

        string password = pwdnode.FirstChild.Value;
        string ipaddress = ipnode.FirstChild.Value;
        // Check to see if the ipaddress belongs to this machine
        if (ipaddresses.Contains(ipaddress))
            return password;
    }
    throw(new SystemException("Did not find password"));
}

Once you have the admin password, you can use it to impersonate the box admin as necessary to run additional code requiring such privileges.   It really helps in allowing us to automatically deploy boxes on GoGrid.  Given the creative commons license of the GoGrid API, the same technique should apply to other cloud providers as necessary.
Hope this helps with your cloud infrastructure deployments – love to hear your comments.

Friday, April 30, 2010

Cheers from the Predixion Dev Team!!

PDT

We’re assembled and ready to rock!  Have a great weekend!

-Jamie and the PX Devs