Saturday, May 05, 2012

EXCLUSIVE: First picture of Barack Obama's first serious girlfriend... and how she went on to marry a Serbian boxer who is now one of New York's most famous bartenders | Mail Online

This is what I was afraid of. Maranis lied about his intentions and now this normal woman is being publicly humiliated and she did not expect what she is getting!!! This is why I lied to the Maranis researcher! 


EXCLUSIVE: First picture of Barack Obama's first serious girlfriend... and how she went on to marry a Serbian boxer who is now one of New York's most famous bartenders

  • Alex McNear met a young Barack Obama at Occidental College in California
  • Two also had a 'summer of romance' in 1981 when Obama transferred to Columbia
  • Once married to Serbian boxer Bob Bozic, who is now a bartender at world-famous Fanelli's Cafe in SoHo
  • Now with child psychologist Robert Stein; they live in Sag Harbor

By Daniel Bates

PUBLISHED: 12:27 EST, 4 May 2012 | UPDATED: 13:09 EST, 4 May 2012

Pictured: Alex McNear also used to date President Obama

Serious girlfriend: Alex McNear was Mr Obama's first serious girlfriend; the two began dating after they met at Occidental College


This is the woman who was Barack Obama's first serious girlfriend, MailOnline can exclusively reveal.

Alex McNear is a mother-of-one who runs a green energy company and lives in a $2million mansion in Sag Harbor, New York.

She is also a Democrat who is understood to have voted for Mr Obama in the 2008 election.

This week it emerged that she won the heart of the future president whilst at college, but until now the story of what happened next to her has not been told.

Rather than moving into the White House she wed a rather less glamorous compensation prize - a Serbian boxer who once tried to rob $60,000 from a bank.

Miss McNear married Bob Bozic in a ceremony which appalled her mother and brought a former vagrant and one time bookmaker into their upper middle class family.

Mr Obama meanwhile romanced his second love Genevieve Cook before meeting Michelle Obama and marrying her - and becoming leader of the Free World.

In her first comments since the story broke, Miss McNear told MailOnline: 'I am not giving any interviews.

'I did that one interview with David Maraniss for reasons I won't go into.

'I have nothing more to say about the whole subject'.

Asked if she would go on a date with Mr Obama if he got back in touch with her, Miss McNear said 'goodbye' and hung up the phone.

Miss McNear was named for the first time this week as the first person Mr Obama fell for in extracts from a forthcoming biopic about the president by author David Maraniss.

She told how the two met at Occidental College in California and had a summer romance at the age of 20 in New York in 1981 when he transferred to Columbia to finish his studies. President Obama's first serious girlfriend, Genevieve Cook, has shared her diaries about their year-long love affair for a new biography.

First love: McNear said she met Mr Obama at Occidental College in California and had a summer romance at the age of 20 in New York in 1981 when he transferred to Columbia to finish his studies

Former flame: Miss McNear, circled, was a literary type from a waspish family who was seduced by Mr Obama's intellectual musings on T.S. Eliot

Former flame: Miss McNear, circled, was a literary type from a waspish family who was seduced by Mr Obama's intellectual musings on T.S. Eliot

The story of what happened next however is equally extraordinary - how a woman who could have been the First Lady ended up with a prize fighter whose claim to fame is losing a 1973 bout with boxer Larry Holmes at Madison Square Garden.

On the face of it, Miss McNear and Mr Bozic made an extremely unlikely pair, even if her mother Suzanne was fiction editor at Playboy in the 1960s and 70s.

Miss McNear was a literary type from a waspish family who was seduced by Mr Obama's intellectual musings on T.S. Eliot.

 

Burly Mr Bozic however spent part of his youth struggling to survive on the streets of Toronto having run away from home.

He was so desperate he would spend his nights sleeping in Laundromats or parked cars and spend his days stealing to makes ends meet.

But love found a way, and in 1987 they got married, and stayed together for seven years, but then divorced.

Bob Bozic - ex-husband to Alex McNear
Bob Bozic - ex-husband to Alex McNear

The boxer: Ms McNear married Bob Bozic in 1987; he famously lost to boxer Larry Holmes at Madison Square Garden in 1973

New vocation: Bozic now works at the world-famous Fanelli Cafe in SoHo

New vocation: Bozic now works at the world-famous Fanelli Cafe in SoHo

They had a daughter together, Vesna, now 20, an aspiring actress who lives in Manhattan and remain close to this day.

In a frank assessment of his function in Miss McNear's life, Mr Bozic has said she married him to 'separate' herself from her family.

And it was not hard to see why.

According to a profile in the New Yorker magazine, he was everything her life was not.

The working class son of an engineer, Mr Bozic got into boxing when the owner of a gym in Toronto took pity on him and gave him a meal every day, making him stay around long enough to get interested in fighting.

In time he became the Canadian national amateur heavyweight champion but went into semi-retirement when he began to lose and travelled around Europe, stopping in Ibiza and Sweden.

His most serious brush with the law came in 1980 when he was broke and in a bad mental place -  supposedly he was sent over the edge when faced with going to the opera on his own.

World-renowned: Fanelli's Cafe in SoHo has been serving drinks since 1922 and is one of New York's oldest continuously-serving bars

World-renowned: Fanelli's Cafe in SoHo has been serving drinks since 1922 and is one of New York's oldest continuously-serving bars

Current partner: McNear is now with Dr Robby Stein

Current partner: McNear is now with Dr Robby Stein, a child psychologist; they live in Sag Harbor

Mr Bozic then walked into a branch of the Manufacturers Hanover bank in Manhattan and told them than men outside had guns pointed at the bank.

He asked for $60,000 and was taken to the vault before the police arrested him. He was given a lenient sentence of probation after admitting robbery in the third degree.

It was after this that he met Miss McNear whilst working as a bouncer at a transvestite club in New York, a job he had only taken to try and keep the judge in his court case happy.

She was working as a lawyer in Chicago at the time phoned up the club asking for a place to hold a benefit night.

He told her his life story and, against all odds, they began to fall for each other.

Happy golden years: Mr Obama is pictured with his grandparents, Stanley and Madelyn Dunham on a park bench outside of New York's Central Park, when Obama was a student at Columbia

Happy golden years: Mr Obama is pictured with his grandparents, Stanley and Madelyn Dunham on a park bench outside of New York's Central Park, when Obama was a student at Columbia

Obama is seen with his father Barack Obama, Sr. in the 1960s
David Maraniss

Family ties: Mr Obama pictured with his father, Barack Obama Sr, in a family photo from the 1960s, left; Mr Obama's ex-girlfriends were interviewed by author David Maraniss, who is releasing a book on the President

Mr Bozic, 61, now works as a barman in New York, a far cry from the affluent lifestyle of his ex-wife.

Miss McNear, 51, runs renewable energy company GreenLogic in Southampton, New York.

Her long term partner is Robert Stein, in his 60s, a child psychologist and their home is on a leafy street in Sag Harbor, which is close the millionaire's playground of The Hamptons.

They two are heavily involved in community affairs and Mr Stein is a member of the village board.

Miss McNear has also worked as a journalist and has written for the local paper the East Hampton Star and attended numerous meetings on zoning issues and developments in the area.

In the extracts of Mr Obama's biography it was his second girlfriend Miss Cook who had harsh words of criticism for him.

But according to the New Yorker piece, Miss McNear also has a sharp tongue when it comes to men. She once said of Mr Bozic: 'Bob is fascinated and appalled by Wasps' - but he still married her.

Speaking from Fanelli's Bar where he works Mr Bozic said: 'I love my ex-wife and I'm very close to her so I'm not going to be able to say anything.'

Asked if he knew the extracts were being made public he said: 'Yes, but she (Miss McNear) didn't know this was going to happen.

First couple: Mr Obama went on to date Genevieve Cook before he eventually met and married Michelle

First couple: Mr Obama went on to date Genevieve Cook before he eventually met and married Michelle

'She didn't know it was going to be like this. She is extremely private and very shy.'

He declined to comment on how he felt about being the man who beat Mr Obama for his ex-wife's affections.

Mr Stein said that he and Miss McNear both wanted to 'put this behind us and get on with our lives'.

Asked whether or not she had told him the extracts about her relationship with Mr Obama were going to be published before they were made public, he said: 'We're very happy for Mr Obama, we're happy he's done so well. I voted for him'.

A friend of Miss McNear added: 'Alex is so embarrassed about this coming out right now. She has a family and is a grown up, for her to be called somebody's 'girlfriend' is really making her cringe.

'It was all in the past but her family still tease her about it. I know her sisters say they would have liked to have been in the White House, but ce la vie.

'To think she could have been First Lady, and she married a boxer instead. You couldn't make it up'.

Another ex: Cook, seen her in college at Swathmore, wrote how Obama loved running and lived in an apartment that smelled of raisins and sweat; she was 25 when they dated
Another ex: Cook, seen her in college at Swathmore, wrote how Obama loved running and lived in an apartment that smelled of raisins and sweat; she was 25 when they dated

Another ex: Cook, seen her at college at Swathmore, wrote how Obama loved running and lived in an apartment that smelled of raisins and sweat; she was 25 when they dated

Love nest: A young Barack Obama lived with his former girlfriend Genevieve Cook on the top floor of this Park Slope townhouse after he graduated from Columbia

Love nest: A young Barack Obama lived with his former girlfriend Genevieve Cook on the top floor of this Park Slope townhouse after he graduated from Columbia

 

Share this article:



Sent from my iPad

Thursday, May 03, 2012

Aurora Fountain Pens

http://www.nibs.com/Auroramainpage.htm


Sent from my iPad

Look at 4 Mexican journalists killed in last week - CBS News

http://www.cbsnews.com/8301-501715_162-57427697/look-at-4-mexican-journalists-killed-in-last-week/?utm_source=feedburner&utm_medium=feed&utm_campaign=Feed%3A+cbsnews%2Ffeed+%28CBSNews.com%29


Sent from my iPad

Daniel Chong, student left handcuffed in DEA cell for 4 days, files $20M claim against agency - CBS News



Sent from my iPad

Seau death: Latest in string of similar deaths of ex-NFL players? - CBS News



Sent from my iPad

Young Barack Obama in Love: A Girlfriend's Secret Diary | Politics | Vanity Fair

http://www.vanityfair.com/politics/2012/06/young-barack-obama-in-love-david-maraniss


Sent from my iPad

Check this out on Remodelista

Rent-a-Garden in Germany
http://remodelista.com/posts/schrebergarten-allotment


Sent from my iPad

Check this out on Remodelista

A Movable Feast: Berlin's Community Garden
http://remodelista.com/posts/a-movable-feast-berlins-community-garden


Sent from my iPad

The best pictures of the Cavalier King

https://www.facebook.com/pages/The-best-pictures-of-the-Cavalier-King/239854999427769


Sent from my iPad

Wednesday, May 02, 2012

Protests Dowtown LA

The protests in downtown Los Angeles seem so anemic. Perhaps because Los Angeles is such a sprawling metropolis that people don't converge in one place, so protests happen at the airport and several other locations, dissipating the impact. And the protests seem rather anarchic with each protestor representing their own personal objections and objectives.

But what really bothers me about these protestors is where have they been? why now? I had to suffer in silence under the years and years of policies that brought our country to this point. It has been a long, long time coming. Why is it only now that there are protests?

I don't know the answers. I can only guess, and my guesses are naturally colored by the lense of my world view. But the truth is that Obama gives people hope that their protests will matter. Living under the Bush/Cheney oligarchic style made people afraid. And that was done deliberately and cynically. Bush/Cheney used the terror warning system to arouse fear in the general population and did not tie the warnings to any actual conditions. Orange did not mean there was knowledge of terrorist activity. Orange meant that Republicans desired to change the conversation and turn it away from real dialogue and back to pedagogy and fear. Nice people.

We should have protested the possibility of invading Iraq. We should have been Legion, in huge multitudes protesting every step of that drum beat to war. We should have been legions outside the United Nations objecting to Colin Powell telling the false stories to defend war.

We should have protested the massive payouts of corporate CEOs at every board meeting. We should have organized as shareholders and pointed out the bad decisions these 'leaders' were being incentivized to make.

We should have protested the corruption on Wall Street long ago. We should have protested the insider trading, stock market manipulation, computerized trading, the accessibility to trade limited to the few. It is beyond frustrating.

For years, I have struggled alone without a voice, dumbfounded by the passivity of the very people being being harmed.

And now finally, there are protests. Where have you been? Where in god's name have you been?

Tuesday, May 01, 2012

A Very Short History of Data Science | What's The Big Data?

A Very Short History of Data Science

I'm in the process of researching the origin and evolution of data science as a discipline and a profession. Here are the milestones that I have picked up so far, tracking the evolution of the term "data science," attempts to define it, and some related developments.  I would greatly appreciate any pointers to additional key milestones (events, publications, etc.).

1974 Peter Naur publishes Concise Survey of Computer Methods in Sweden and the United States. The book is a survey of contemporary data processing methods that are used in a wide range of applications. It is organized around the concept of data as defined in the IFIP Guide to Concepts and Terms in Data Processing, which defines data as "a representation of facts or ideas in a formalized manner capable of being communicated or manipulated by some process." The Preface to the book tells the reader that a course plan was presented at the IFIP Congress in 1968, titled "Datalogy, the science of data and of data processes and its place in education," and that in the text of the book, "the term 'data science' has been used freely." Naur offers the following definition of data science: "The science of dealing with data, once they have been established, while the relation of the data to what they represent is delegated to other fields and sciences."

1977 The International Association for Statistical Computing (IASC) was founded as a Section of the ISI. "It is the mission of the IASC to link traditional statistical methodology, modern computer technology, and the knowledge of domain experts in order to convert data into information and knowledge."

1989 Gregory Piatetsky-Shapiro organizes and chairs the first Knowledge Discovery in Databases (KDD) workshop. In 1995, it became the annual ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD).

1996 Members of the International Federation of Classification Societies (IFCS) meet in Tokyo for their biennial conference. For the first time, the term "data science" is included in the title of the conference ("Data science, classification, and related methods"). The IFCS was founded in 1985 by six country- and language-specific classification societies, one of which, The Classification Society, was founded in 1964. The aim of these classification societies has been to support the study of "the principle and practice of classification in a wide range of disciplines"(CS), "research in problems of classification, data analysis, and systems for ordering knowledge"(IFCS), and the "study of classification and clustering (including systematic methods of creating classifications from data) and related statistical and data analytic methods" (CSNA bylaws). The classification societies have variously used the terms data analysis, data mining, and data science in their publications.

1997 Launch of the journal Knowledge Discovery and Data Mining: "Advances in data gathering, storage, and distribution have created a need for computational tools and techniques to aid in data analysis. Data Mining and Knowledge Discovery in Databases (KDD) is a rapidly growing area of research and application that builds on techniques and theories from many fields, including statistics, databases, pattern recognition and learning, data visualization, uncertainty modelling, data warehousing and OLAP, optimization, and high performance computing. KDD is concerned with issues of scalability, the multi-step knowledge discovery process for extracting useful patterns and models from raw data stores (including data cleaning and noise modelling), and issues of making discovered patterns understandable."

2001 William S. Cleveland  (then at Bell Labs, now at the Department of Statistics at Purdue University) publishes "Data Science: An Action Plan for Expanding the Technical Areas of the Field of Statistics." It is a plan "to enlarge the major areas of technical work of the field of statistics. Because the plan is ambitious and implies substantial change, the altered field will be called 'data science.'" The plan "sets out six technical areas for a university department": Multidisciplinary Investigations, Models and Methods for Data, Computing with Data, Pedagogy, Tool Evaluation, and Theory. Cleveland puts the proposed new discipline in the context of computer science and the contemporary work on data mining: "…the benefit to the data analyst has been limited, because the knowledge among computer scientists about how to think of and approach the analysis of data is limited, just as the knowledge of computing environments by statisticians is limited. A merger of knowledge bases would produce a powerful force for innovation. This suggests that statisticians should look to computing for knowledge today just as data science looked to mathematics in the past. … departments of data science should contain faculty members who devote their careers to advances in computing with data and who form partnership with computer scientists."

April 2002 The Data Science Journal is launched, publishing papers on "the management of data and databases in Science and Technology. The scope of the Journal includes descriptions of data systems, their publication on the internet, applications and legal issues." The journal is published by the Committee on Data for Science and Technology (CODATA) of the International Council for Science (ICSU).

January 2003 The Journal of Data Science is launched: "By 'Data Science' we mean almost everything that has something to do with data: Collecting, analyzing, modeling…… yet the most important part is its applications — all sorts of applications. This journal is devoted to applications of statistical methods at large…. The Journal of Data Science will provide a platform for all data workers to present their views and exchange ideas."

September 2005 The National Science Board publishes "Long-lived Digital Data Collections: Enabling Research and Education in the 21st Century." One of the recommendations of the report reads: "The NSF, working in partnership with collection managers and the community at large, should act to develop and mature the career path for data scientists and to ensure that the research enterprise includes a sufficient number of high-quality data scientists." The report defines data scientists as "the information and computer scientists, database and software engineers and programmers, disciplinary experts, curators and expert annotators, librarians, archivists, and others, who are crucial to the successful management of a digital data collection."

July 2008 The JISC publishes the final report of a study it commissioned to "examine and make recommendations on the role and career development of data scientists and the associated supply of specialist data curation skills to the research community. " The study's final report, "The Skills, Role & Career Structure of Data Scientists & Curators:  Assessment of Current Practice & Future Needs," defines data scientists as "people who work where the research is carried out – or, in the case of data centre personnel, in close collaboration with the creators of the data – and may be involved in creative enquiry and analysis, enabling others to work with digital data, and developments in data base technology."

January 2009 Harnessing the Power of Digital Data for Science and Society is published. This report of the Interagency Working Group on Digital Data to the Committee on Science of the National Science and Technology Council states that "The nation needs to identify and promote the emergence of new disciplines and specialists expert in addressing the complex and dynamic challenges of digital preservation, sustained access, reuse and repurposing of data. Many disciplines are seeing the emergence of a new type of data science and management expert, accomplished in the computer, information, and data sciences arenas and in another domain science. These individuals are key to the current and future success of the scientific enterprise. However, these individuals often receive little recognition for their contributions and have limited career paths. Critical challenges in achieving our strategic vision include providing an effective pipeline of data professionals to ensure that the needs and opportunities of the future can be met and providing these professionals with appropriate rewards and recognition." The report discusses the emergence of "new information disciplines" and lists a few examples:

  • Digital Curators: experts knowledgeable of and with responsibility for the content of digital collection(s);
  • Digital Archivists: experts competent to appraise, acquire, authenticate, preserve, and provide access to records in digital form; and
  • Data Scientists: information and computer scientists, database and software engineers and programmers, disciplinary experts, expert annotators, and others who are crucial to the successful management of a digital data collection.

May 2009 Mike Driscoll writes in "The Three Sexy Skills of Data Geeks": "…with the Age of Data upon us, those who can model, munge, and visually communicate data — call us statisticians or data geeks — are a hot commodity." [Driscoll will follow up with The Seven Secrets of Successful Data Scientists in August 2010]

June 2009 Nathan Yau writes in "Rise of the Data Scientist":  "As we've all read by now, Google's chief economist Hal Varian commented in January that the next sexy job in the next 10 years would be statisticians. Obviously, I whole-heartedly agree. Heck, I'd go a step further and say they're sexy now – mentally and physically. However, if you went on to read the rest of Varian's interview, you'd know that by statisticians, he actually meant it as a general title for someone who is able to extract information from large datasets and then present something of use to non-data experts… [Ben] Fry… argues for an entirely new field that combines the skills and talents from often disjoint areas of expertise… [computer science; mathematics, statistics, and data mining; graphic design; infovis and human-computer interaction]. And after two years of highlighting visualization on FlowingData, it seems collaborations between the fields are growing more common, but more importantly, computational information design edges closer to reality. We're seeing data scientists – people who can do it all – emerge from the rest of the pack."

June 2009 Troy Sadkowsky creates the data scientists group on LinkedIn as a companion to his website, datasceintists.com (which later became datascientists.net).

[update]February 2010 Kenneth Cukier writes in "Data, data everywhere: A special report on managing information": "… a new kind of professional has emerged, the data scientist, who combines the skills of software programmer, statistician and storyteller/artist to extract the nuggets of gold hidden under mountains of data."

June 2010 Mike Loukides writes in "What is Data Science?":  "Data scientists combine entrepreneurship with patience, the willingness to build data products incrementally, the ability to explore, and the ability to iterate over a solution. They are inherently interdisciplinary. They can tackle all aspects of a problem, from initial data collection and data conditioning to drawing conclusions. They can think outside the box to come up with new ways to view the problem, or to work with very broadly defined problems: 'here's a lot of data, what can you make from it?'"

September 2010  Hilary Mason and Chris Wiggins write in "A Taxonomy of Data Science":  "…we thought it would be useful to propose one possible taxonomy… of what a data scientist does, in roughly chronological order: Obtain, Scrub, Explore, Model, and iNterpret…. Data science is clearly a blend of the hackers' arts… statistics and machine learning… and the expertise in mathematics and the domain of the data for the analysis to be interpretable… It requires creative decisions and open-mindedness in a scientific context."

September 2010 Drew Conway writes in "The Data Science Venn Diagram":  "…one needs to learn a lot as they aspire to become a fully competent data scientist. Unfortunately, simply enumerating texts and tutorials does not untangle the knots. Therefore, in an effort to simplify the discussion, and add my own thoughts to what is already a crowded market of ideas, I present the Data Science Venn Diagram… hacking skills, math and stats knowledge, and substantive expertise."

May 2011  Pete Warden writes in "Why the term 'data science' is flawed but useful": "There is no widely accepted boundary for what's inside and outside of data science's scope. Is it just a faddish rebranding of statistics? I don't think so, but I also don't have a full definition. I believe that the recent abundance of data has sparked something new in the world, and when I look around I see people with shared characteristics who don't fit into traditional categories. These people tend to work beyond the narrow specialties that dominate the corporate and institutional world, handling everything from finding the data, processing it at scale, visualizing it and writing it up as a story. They also seem to start by looking at what the data can tell them, and then picking interesting threads to follow, rather than the traditional scientist's approach of choosing the problem first and then finding data to shed light on it."

May 2011 David Smith writes in "'Data Science': What's in a name?":   "The terms 'Data Science' and 'Data Scientist' have only been in common usage for a little over a year, but they've really taken off since then: many companies are now hiring for 'data scientists', and entire conferences are run under the name of 'data science'. But despite the widespread adoption, some have resisted the change from the more traditional terms like 'statistician' or 'quant' or 'data analyst'…. I think 'Data Science' better describes what we actually do: a combination of computer hacking, data analysis, and problem solving."

September 2011 Harlan Harris writes in "Data Science, Moore's Law, and Moneyball" : "'Data Science' is defined as what 'Data Scientists' do. What Data Scientists do has been very well covered, and it runs the gamut from data collection and munging, through application of statistics and machine learning and related techniques, to interpretation, communication, and visualization of the results. Who Data Scientists are may be the more fundamental question…  I tend to like the idea that Data Science is defined by its practitioners, that it's a career path rather than a category of activities. In my conversations with people, it seems that people who consider themselves Data Scientists typically have eclectic career paths, that might in some ways seem not to make much sense."

September 2011 DJ Patil writes in "Building Data Science Teams": "Starting in 2008, Jeff Hammerbacher (@hackingdata) and I sat down to share our experiences building the data and analytics groups at Facebook and LinkedIn. In many ways, that meeting was the start of data science as a distinct professional specialization….  we realized that as our organizations grew, we both had to figure out what to call the people on our teams. 'Business analyst' seemed too limiting. 'Data analyst' was a contender, but we felt that title might limit what people could do. After all, many of the people on our teams had deep engineering expertise. 'Research scientist' was a reasonable job title used by companies like Sun, HP, Xerox, Yahoo, and IBM. However, we felt that most research scientists worked on projects that were futuristic and abstract, and the work was done in labs that were isolated from the product development teams. It might take years for lab research to affect key products, if it ever did. Instead, the focus of our teams was to work on data applications that would have an immediate and massive impact on the business. The term that seemed to fit best was data scientist: those who use both data and science to create something new. "

Update: Gregory Piatetsky-Shapiro posted a great discussion of the journey from data mining to big data. Note that in this timeline, I tried to focus on specific mentions of "data science" and attempts to define it.

I'm in the process of researching the origin and evolution of data science as a discipline and a profession. Here are the milestones that I have picked up so far, tracking the evolution of the term "data science," attempts to define it, and some related developments.  I would greatly appreciate any pointers to additional key milestones (events, publications, etc.).

1974 Peter Naur publishes Concise Survey of Computer Methods in Sweden and the United States. The book is a survey of contemporary data processing methods that are used in a wide range of applications. It is organized around the concept of data as defined in the IFIP Guide to Concepts and Terms in Data Processing, which defines data as "a representation of facts or ideas in a formalized manner capable of being communicated or manipulated by some process." The Preface to the book tells the reader that a course plan was presented at the IFIP Congress in 1968, titled "Datalogy, the science of data and of data processes and its place in education," and that in the text of the book, "the term 'data science' has been used freely." Naur offers the following definition of data science: "The science of dealing with data, once they have been established, while the relation of the data to what they represent is delegated to other fields and sciences."

1977 The International Association for Statistical Computing (IASC) was founded as a Section of the ISI. "It is the mission of the IASC to link traditional statistical methodology, modern computer technology, and the knowledge of domain experts in order to convert data into information and knowledge."

1989 Gregory Piatetsky-Shapiro organizes and chairs the first Knowledge Discovery in Databases (KDD) workshop. In 1995, it became the annual ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD).

1996 Members of the International Federation of Classification Societies (IFCS) meet in Tokyo for their biennial conference. For the first time, the term "data science" is included in the title of the conference ("Data science, classification, and related methods"). The IFCS was founded in 1985 by six country- and language-specific classification societies, one of which, The Classification Society, was founded in 1964. The aim of these classification societies has been to support the study of "the principle and practice of classification in a wide range of disciplines"(CS), "research in problems of classification, data analysis, and systems for ordering knowledge"(IFCS), and the "study of classification and clustering (including systematic methods of creating classifications from data) and related statistical and data analytic methods" (CSNA bylaws). The classification societies have variously used the terms data analysis, data mining, and data science in their publications.

1997 Launch of the journal Knowledge Discovery and Data Mining: "Advances in data gathering, storage, and distribution have created a need for computational tools and techniques to aid in data analysis. Data Mining and Knowledge Discovery in Databases (KDD) is a rapidly growing area of research and application that builds on techniques and theories from many fields, including statistics, databases, pattern recognition and learning, data visualization, uncertainty modelling, data warehousing and OLAP, optimization, and high performance computing. KDD is concerned with issues of scalability, the multi-step knowledge discovery process for extracting useful patterns and models from raw data stores (including data cleaning and noise modelling), and issues of making discovered patterns understandable."

2001 William S. Cleveland  (then at Bell Labs, now at the Department of Statistics at Purdue University) publishes "Data Science: An Action Plan for Expanding the Technical Areas of the Field of Statistics." It is a plan "to enlarge the major areas of technical work of the field of statistics. Because the plan is ambitious and implies substantial change, the altered field will be called 'data science.'" The plan "sets out six technical areas for a university department": Multidisciplinary Investigations, Models and Methods for Data, Computing with Data, Pedagogy, Tool Evaluation, and Theory. Cleveland puts the proposed new discipline in the context of computer science and the contemporary work on data mining: "…the benefit to the data analyst has been limited, because the knowledge among computer scientists about how to think of and approach the analysis of data is limited, just as the knowledge of computing environments by statisticians is limited. A merger of knowledge bases would produce a powerful force for innovation. This suggests that statisticians should look to computing for knowledge today just as data science looked to mathematics in the past. … departments of data science should contain faculty members who devote their careers to advances in computing with data and who form partnership with computer scientists."

April 2002 The Data Science Journal is launched, publishing papers on "the management of data and databases in Science and Technology. The scope of the Journal includes descriptions of data systems, their publication on the internet, applications and legal issues." The journal is published by the Committee on Data for Science and Technology (CODATA) of the International Council for Science (ICSU).

January 2003 The Journal of Data Science is launched: "By 'Data Science' we mean almost everything that has something to do with data: Collecting, analyzing, modeling…… yet the most important part is its applications — all sorts of applications. This journal is devoted to applications of statistical methods at large…. The Journal of Data Science will provide a platform for all data workers to present their views and exchange ideas."

September 2005 The National Science Board publishes "Long-lived Digital Data Collections: Enabling Research and Education in the 21st Century." One of the recommendations of the report reads: "The NSF, working in partnership with collection managers and the community at large, should act to develop and mature the career path for data scientists and to ensure that the research enterprise includes a sufficient number of high-quality data scientists." The report defines data scientists as "the information and computer scientists, database and software engineers and programmers, disciplinary experts, curators and expert annotators, librarians, archivists, and others, who are crucial to the successful management of a digital data collection."

July 2008 The JISC publishes the final report of a study it commissioned to "examine and make recommendations on the role and career development of data scientists and the associated supply of specialist data curation skills to the research community. " The study's final report, "The Skills, Role & Career Structure of Data Scientists & Curators:  Assessment of Current Practice & Future Needs," defines data scientists as "people who work where the research is carried out – or, in the case of data centre personnel, in close collaboration with the creators of the data – and may be involved in creative enquiry and analysis, enabling others to work with digital data, and developments in data base technology."

January 2009 Harnessing the Power of Digital Data for Science and Society is published. This report of the Interagency Working Group on Digital Data to the Committee on Science of the National Science and Technology Council states that "The nation needs to identify and promote the emergence of new disciplines and specialists expert in addressing the complex and dynamic challenges of digital preservation, sustained access, reuse and repurposing of data. Many disciplines are seeing the emergence of a new type of data science and management expert, accomplished in the computer, information, and data sciences arenas and in another domain science. These individuals are key to the current and future success of the scientific enterprise. However, these individuals often receive little recognition for their contributions and have limited career paths. Critical challenges in achieving our strategic vision include providing an effective pipeline of data professionals to ensure that the needs and opportunities of the future can be met and providing these professionals with appropriate rewards and recognition." The report discusses the emergence of "new information disciplines" and lists a few examples:

  • Digital Curators: experts knowledgeable of and with responsibility for the content of digital collection(s);
  • Digital Archivists: experts competent to appraise, acquire, authenticate, preserve, and provide access to records in digital form; and
  • Data Scientists: information and computer scientists, database and software engineers and programmers, disciplinary experts, expert annotators, and others who are crucial to the successful management of a digital data collection.

May 2009 Mike Driscoll writes in "The Three Sexy Skills of Data Geeks": "…with the Age of Data upon us, those who can model, munge, and visually communicate data — call us statisticians or data geeks — are a hot commodity." [Driscoll will follow up with The Seven Secrets of Successful Data Scientists in August 2010]

June 2009 Nathan Yau writes in "Rise of the Data Scientist":  "As we've all read by now, Google's chief economist Hal Varian commented in January that the next sexy job in the next 10 years would be statisticians. Obviously, I whole-heartedly agree. Heck, I'd go a step further and say they're sexy now – mentally and physically. However, if you went on to read the rest of Varian's interview, you'd know that by statisticians, he actually meant it as a general title for someone who is able to extract information from large datasets and then present something of use to non-data experts… [Ben] Fry… argues for an entirely new field that combines the skills and talents from often disjoint areas of expertise… [computer science; mathematics, statistics, and data mining; graphic design; infovis and human-computer interaction]. And after two years of highlighting visualization on FlowingData, it seems collaborations between the fields are growing more common, but more importantly, computational information design edges closer to reality. We're seeing data scientists – people who can do it all – emerge from the rest of the pack."

June 2009 Troy Sadkowsky creates the data scientists group on LinkedIn as a companion to his website, datasceintists.com (which later became datascientists.net).

[update]February 2010 Kenneth Cukier writes in "Data, data everywhere: A special report on managing information": "… a new kind of professional has emerged, the data scientist, who combines the skills of software programmer, statistician and storyteller/artist to extract the nuggets of gold hidden under mountains of data."

June 2010 Mike Loukides writes in "What is Data Science?":  "Data scientists combine entrepreneurship with patience, the willingness to build data products incrementally, the ability to explore, and the ability to iterate over a solution. They are inherently interdisciplinary. They can tackle all aspects of a problem, from initial data collection and data conditioning to drawing conclusions. They can think outside the box to come up with new ways to view the problem, or to work with very broadly defined problems: 'here's a lot of data, what can you make from it?'"

September 2010  Hilary Mason and Chris Wiggins write in "A Taxonomy of Data Science":  "…we thought it would be useful to propose one possible taxonomy… of what a data scientist does, in roughly chronological order: Obtain, Scrub, Explore, Model, and iNterpret…. Data science is clearly a blend of the hackers' arts… statistics and machine learning… and the expertise in mathematics and the domain of the data for the analysis to be interpretable… It requires creative decisions and open-mindedness in a scientific context."

September 2010 Drew Conway writes in "The Data Science Venn Diagram":  "…one needs to learn a lot as they aspire to become a fully competent data scientist. Unfortunately, simply enumerating texts and tutorials does not untangle the knots. Therefore, in an effort to simplify the discussion, and add my own thoughts to what is already a crowded market of ideas, I present the Data Science Venn Diagram… hacking skills, math and stats knowledge, and substantive expertise."

May 2011  Pete Warden writes in "Why the term 'data science' is flawed but useful": "There is no widely accepted boundary for what's inside and outside of data science's scope. Is it just a faddish rebranding of statistics? I don't think so, but I also don't have a full definition. I believe that the recent abundance of data has sparked something new in the world, and when I look around I see people with shared characteristics who don't fit into traditional categories. These people tend to work beyond the narrow specialties that dominate the corporate and institutional world, handling everything from finding the data, processing it at scale, visualizing it and writing it up as a story. They also seem to start by looking at what the data can tell them, and then picking interesting threads to follow, rather than the traditional scientist's approach of choosing the problem first and then finding data to shed light on it."

May 2011 David Smith writes in "'Data Science': What's in a name?":   "The terms 'Data Science' and 'Data Scientist' have only been in common usage for a little over a year, but they've really taken off since then: many companies are now hiring for 'data scientists', and entire conferences are run under the name of 'data science'. But despite the widespread adoption, some have resisted the change from the more traditional terms like 'statistician' or 'quant' or 'data analyst'…. I think 'Data Science' better describes what we actually do: a combination of computer hacking, data analysis, and problem solving."

September 2011 Harlan Harris writes in "Data Science, Moore's Law, and Moneyball" : "'Data Science' is defined as what 'Data Scientists' do. What Data Scientists do has been very well covered, and it runs the gamut from data collection and munging, through application of statistics and machine learning and related techniques, to interpretation, communication, and visualization of the results. Who Data Scientists are may be the more fundamental question…  I tend to like the idea that Data Science is defined by its practitioners, that it's a career path rather than a category of activities. In my conversations with people, it seems that people who consider themselves Data Scientists typically have eclectic career paths, that might in some ways seem not to make much sense."

September 2011 DJ Patil writes in "Building Data Science Teams": "Starting in 2008, Jeff Hammerbacher (@hackingdata) and I sat down to share our experiences building the data and analytics groups at Facebook and LinkedIn. In many ways, that meeting was the start of data science as a distinct professional specialization….  we realized that as our organizations grew, we both had to figure out what to call the people on our teams. 'Business analyst' seemed too limiting. 'Data analyst' was a contender, but we felt that title might limit what people could do. After all, many of the people on our teams had deep engineering expertise. 'Research scientist' was a reasonable job title used by companies like Sun, HP, Xerox, Yahoo, and IBM. However, we felt that most research scientists worked on projects that were futuristic and abstract, and the work was done in labs that were isolated from the product development teams. It might take years for lab research to affect key products, if it ever did. Instead, the focus of our teams was to work on data applications that would have an immediate and massive impact on the business. The term that seemed to fit best was data scientist: those who use both data and science to create something new. "

Update: Gregory Piatetsky-Shapiro posted a great discussion of the journey from data mining to big data. Note that in this timeline, I tried to focus on specific mentions of "data science" and attempts to define it.



Sent from my iPad

Zen Moments Bookstore - Make Me an Instrument of Your Peace

Product Description

"The Cab Ride I'll Never Forget" appears in this beautiful book - which takes the prayer of St Francis into everyday life.

Kent Nerburn's Make Me an Instrument of Your Peace, immerses us in the spirit of one of the most universally inspiring figures in history: St. Francis of Assisi. The Prayer of St. Francis boldly but gently challenges us to resist the forces of evil and negativity with the spirit of goodwill and generosity. And Nerburn shows, in his wonderfully personal and humble way, how we each can live out the prayer's prescription for living in our everyday and less-than-saintly lives.

"Where there is hatred, let me sow love...Where there is injury, let me sow pardon..." Expanding upon each line of the St. Francis Prayer, Nerburn shares touching, inspiring stories from his own experience and that of others and reveals how each of us can make a difference for good in ordinary ways without being heroes or saints. Struggling to help a young son comfort his best friend when his mother dies, moved by the courage of war enemies who reconcile, being wrenched out of self-absorbed depression by responding to someone else's tragedy, taking a spirited old lady on a farewell taxi ride through her town-these are the kinds of everyday moments in which Nerburn finds we can live out the spirit of St. Francis.

By incorporating the power and grace of these few lines of practical idealism into our thoughts and deeds, we can begin to ease our own suffering-and the suffering of those with whom we share our lives. And, remarkably, find a way to true peace and happiness by tapping into our basic human goodness. As we open our hearts and embrace his words, St. Francis "touches our deepest humanity and ignites the spark of our divinity."

Lord, make me an instrument of thy peace.
Where there is hatred let me sow love,
Where there is injury let me sow pardon,
Where there is doubt, faith,
Where there is despair, hope,
Where there is darkness, light,
And where there is sadness, joy...

In this beautifully written book, Kent Nerburn leads us into the heart of the St. Francis Prayer and line by line demonstrates how St. Francis's words can resonate in our lives today.



Sent from my iPad

Bill Text - SB-1520 Force fed birds.

FINALLY!!! Only took 8 years:

LEGISLATIVE COUNSEL'S DIGEST

[ Filed Secretary of State  September 29, 2004. Approved by Governor  September 29, 2004. ]

Existing law authorizes an officer to issue a citation to a person or entity keeping horses or other equine animals for hire if the person or entity fails to meet standards of humane treatment regarding the keeping of horses or other equine animals.

This bill would establish similar provisions regarding force feeding a bird, as defined. The bill would prohibit a person from force feeding a bird for the purpose of enlarging the bird's liver beyond normal size, and would prohibit a person from hiring another person to do so. The bill would also prohibit a product from being sold in the state if it is the result of force feeding a bird for the purpose of enlarging the bird's liver beyond normal size. The bill would authorize an officer to issue a citation for a violation of those provisions in an amount up to $1,000 per violation per day.

The bill would provide that these prohibitions shall become operative on July 1, 2012.

Until July 1, 2012, this bill would prohibit an existing or future civil or criminal cause of action for engaging in an act prohibited by the bill, from proceeding against a person or entity engaged in, or controlled by persons or entities who were engaged in, agricultural practices that involved force feeding birds at the time of the enactment of this bill.



Sent from my iPad

What is data science? - O'Reilly Radar

What is data science?

Report sections

What is Data Science?

Download free PDF version

We've all heard it: according to Hal Varian, statistics is the next sexy job. Five years ago, in What is Web 2.0, Tim O'Reilly said that "data is the next Intel Inside." But what does that statement mean? Why do we suddenly care about statistics and about data?

In this post, I examine the many sides of data science -- the technologies, the companies and the unique skill sets.

The web is full of "data-driven apps." Almost any e-commerce application is a data-driven application. There's a database behind a web front end, and middleware that talks to a number of other databases and data services (credit card processing companies, banks, and so on). But merely using data isn't really what we mean by "data science." A data application acquires its value from the data itself, and creates more data as a result. It's not just an application with data; it's a data product. Data science enables the creation of data products.

One of the earlier data products on the Web was the CDDB database. The developers of CDDB realized that any CD had a unique signature, based on the exact length (in samples) of each track on the CD. Gracenote built a database of track lengths, and coupled it to a database of album metadata (track titles, artists, album titles). If you've ever used iTunes to rip a CD, you've taken advantage of this database. Before it does anything else, iTunes reads the length of every track, sends it to CDDB, and gets back the track titles. If you have a CD that's not in the database (including a CD you've made yourself), you can create an entry for an unknown album. While this sounds simple enough, it's revolutionary: CDDB views music as data, not as audio, and creates new value in doing so. Their business is fundamentally different from selling music, sharing music, or analyzing musical tastes (though these can also be "data products"). CDDB arises entirely from viewing a musical problem as a data problem.

Google is a master at creating data products. Here's a few examples:

  • Google's breakthrough was realizing that a search engine could use input other than the text on the page. Google's PageRank algorithm was among the first to use data outside of the page itself, in particular, the number of links pointing to a page. Tracking links made Google searches much more useful, and PageRank has been a key ingredient to the company's success.
  • Spell checking isn't a terribly difficult problem, but by suggesting corrections to misspelled searches, and observing what the user clicks in response, Google made it much more accurate. They've built a dictionary of common misspellings, their corrections, and the contexts in which they occur.
  • Speech recognition has always been a hard problem, and it remains difficult. But Google has made huge strides by using the voice data they've collected, and has been able to integrate voice search into their core search engine.
  • During the Swine Flu epidemic of 2009, Google was able to track the progress of the epidemic by following searches for flu-related topics.

Flu trends

Google was able to spot trends in the Swine Flu epidemic roughly two weeks before the Center for Disease Control by analyzing searches that people were making in different regions of the country.

Google isn't the only company that knows how to use data. Facebook and LinkedIn use patterns of friendship relationships to suggest other people you may know, or should know, with sometimes frightening accuracy. Amazon saves your searches, correlates what you search for with what other users search for, and uses it to create surprisingly appropriate recommendations. These recommendations are "data products" that help to drive Amazon's more traditional retail business. They come about because Amazon understands that a book isn't just a book, a camera isn't just a camera, and a customer isn't just a customer; customers generate a trail of "data exhaust" that can be mined and put to use, and a camera is a cloud of data that can be correlated with the customers' behavior, the data they leave every time they visit the site.

The thread that ties most of these applications together is that data collected from users provides added value. Whether that data is search terms, voice samples, or product reviews, the users are in a feedback loop in which they contribute to the products they use. That's the beginning of data science.

In the last few years, there has been an explosion in the amount of data that's available. Whether we're talking about web server logs, tweet streams, online transaction records, "citizen science," data from sensors, government data, or some other source, the problem isn't finding data, it's figuring out what to do with it. And it's not just companies using their own data, or the data contributed by their users. It's increasingly common to mashup data from a number of sources. "Data Mashups in R" analyzes mortgage foreclosures in Philadelphia County by taking a public report from the county sheriff's office, extracting addresses and using Yahoo to convert the addresses to latitude and longitude, then using the geographical data to place the foreclosures on a map (another data source), and group them by neighborhood, valuation, neighborhood per-capita income, and other socio-economic factors.

The question facing every company today, every startup, every non-profit, every project site that wants to attract a community, is how to use data effectively -- not just their own data, but all the data that's available and relevant. Using data effectively requires something different from traditional statistics, where actuaries in business suits perform arcane but fairly well-defined kinds of analysis. What differentiates data science from statistics is that data science is a holistic approach. We're increasingly finding data in the wild, and data scientists are involved with gathering data, massaging it into a tractable form, making it tell its story, and presenting that story to others.

To get a sense for what skills are required, let's look at the data lifecycle: where it comes from, how you use it, and where it goes.

Where data comes from

Data is everywhere: your government, your web server, your business partners, even your body. While we aren't drowning in a sea of data, we're finding that almost everything can (or has) been instrumented. At O'Reilly, we frequently combine publishing industry data from Nielsen BookScan with our own sales data, publicly available Amazon data, and even job data to see what's happening in the publishing industry. Sites like Infochimps and Factual provide access to many large datasets, including climate data, MySpace activity streams, and game logs from sporting events. Factual enlists users to update and improve its datasets, which cover topics as diverse as endocrinologists to hiking trails.

1956 disk drive

One of the first commercial disk drives from IBM. It has a 5 MB capacity and it's stored in a cabinet roughly the size of a luxury refrigerator. In contrast, a 32 GB microSD card measures around 5/8 x 3/8 inch and weighs about 0.5 gram.

Photo: Mike Loukides. Disk drive on display at IBM Almaden Research

Much of the data we currently work with is the direct consequence of Web 2.0, and of Moore's Law applied to data. The web has people spending more time online, and leaving a trail of data wherever they go. Mobile applications leave an even richer data trail, since many of them are annotated with geolocation, or involve video or audio, all of which can be mined. Point-of-sale devices and frequent-shopper's cards make it possible to capture all of your retail transactions, not just the ones you make online. All of this data would be useless if we couldn't store it, and that's where Moore's Law comes in. Since the early '80s, processor speed has increased from 10 MHz to 3.6 GHz -- an increase of 360 (not counting increases in word length and number of cores). But we've seen much bigger increases in storage capacity, on every level. RAM has moved from $1,000/MB to roughly $25/GB -- a price reduction of about 40000, to say nothing of the reduction in size and increase in speed. Hitachi made the first gigabyte disk drives in 1982, weighing in at roughly 250 pounds; now terabyte drives are consumer equipment, and a 32 GB microSD card weighs about half a gram. Whether you look at bits per gram, bits per dollar, or raw capacity, storage has more than kept pace with the increase of CPU speed.

The importance of Moore's law as applied to data isn't just geek pyrotechnics. Data expands to fill the space you have to store it. The more storage is available, the more data you will find to put into it. The data exhaust you leave behind whenever you surf the web, friend someone on Facebook, or make a purchase in your local supermarket, is all carefully collected and analyzed. Increased storage capacity demands increased sophistication in the analysis and use of that data. That's the foundation of data science.

So, how do we make that data useful? The first step of any data analysis project is "data conditioning," or getting data into a state where it's usable. We are seeing more data in formats that are easier to consume: Atom data feeds, web services, microformats, and other newer technologies provide data in formats that's directly machine-consumable. But old-style screen scraping hasn't died, and isn't going to die. Many sources of "wild data" are extremely messy. They aren't well-behaved XML files with all the metadata nicely in place. The foreclosure data used in "Data Mashups in R" was posted on a public website by the Philadelphia county sheriff's office. This data was presented as an HTML file that was probably generated automatically from a spreadsheet. If you've ever seen the HTML that's generated by Excel, you know that's going to be fun to process.

Data conditioning can involve cleaning up messy HTML with tools like Beautiful Soup, natural language processing to parse plain text in English and other languages, or even getting humans to do the dirty work. You're likely to be dealing with an array of data sources, all in different forms. It would be nice if there was a standard set of tools to do the job, but there isn't. To do data conditioning, you have to be ready for whatever comes, and be willing to use anything from ancient Unix utilities such as awk to XML parsers and machine learning libraries. Scripting languages, such as Perl and Python, are essential.

Once you've parsed the data, you can start thinking about the quality of your data. Data is frequently missing or incongruous. If data is missing, do you simply ignore the missing points? That isn't always possible. If data is incongruous, do you decide that something is wrong with badly behaved data (after all, equipment fails), or that the incongruous data is telling its own story, which may be more interesting? It's reported that the discovery of ozone layer depletion was delayed because automated data collection tools discarded readings that were too low . In data science, what you have is frequently all you're going to get. It's usually impossible to get "better" data, and you have no alternative but to work with the data at hand.

If the problem involves human language, understanding the data adds another dimension to the problem. Roger Magoulas, who runs the data analysis group at O'Reilly, was recently searching a database for Apple job listings requiring geolocation skills. While that sounds like a simple task, the trick was disambiguating "Apple" from many job postings in the growing Apple industry. To do it well you need to understand the grammatical structure of a job posting; you need to be able to parse the English. And that problem is showing up more and more frequently. Try using Google Trends to figure out what's happening with the Cassandra database or the Python language, and you'll get a sense of the problem. Google has indexed many, many websites about large snakes. Disambiguation is never an easy task, but tools like the Natural Language Toolkit library can make it simpler.

When natural language processing fails, you can replace artificial intelligence with human intelligence. That's where services like Amazon's Mechanical Turk come in. If you can split your task up into a large number of subtasks that are easily described, you can use Mechanical Turk's marketplace for cheap labor. For example, if you're looking at job listings, and want to know which originated with Apple, you can have real people do the classification for roughly $0.01 each. If you have already reduced the set to 10,000 postings with the word "Apple," paying humans $0.01 to classify them only costs $100.

Working with data at scale

We've all heard a lot about "big data," but "big" is really a red herring. Oil companies, telecommunications companies, and other data-centric industries have had huge datasets for a long time. And as storage capacity continues to expand, today's "big" is certainly tomorrow's "medium" and next week's "small." The most meaningful definition I've heard: "big data" is when the size of the data itself becomes part of the problem. We're discussing data problems ranging from gigabytes to petabytes of data. At some point, traditional techniques for working with data run out of steam.

What are we trying to do with data that's different? According to Jeff Hammerbacher (@hackingdata), we're trying to build information platforms or dataspaces. Information platforms are similar to traditional data warehouses, but different. They expose rich APIs, and are designed for exploring and understanding the data rather than for traditional analysis and reporting. They accept all data formats, including the most messy, and their schemas evolve as the understanding of the data changes.

Most of the organizations that have built data platforms have found it necessary to go beyond the relational database model. Traditional relational database systems stop being effective at this scale. Managing sharding and replication across a horde of database servers is difficult and slow. The need to define a schema in advance conflicts with reality of multiple, unstructured data sources, in which you may not know what's important until after you've analyzed the data. Relational databases are designed for consistency, to support complex transactions that can easily be rolled back if any one of a complex set of operations fails. While rock-solid consistency is crucial to many applications, it's not really necessary for the kind of analysis we're discussing here. Do you really care if you have 1,010 or 1,012 Twitter followers? Precision has an allure, but in most data-driven applications outside of finance, that allure is deceptive. Most data analysis is comparative: if you're asking whether sales to Northern Europe are increasing faster than sales to Southern Europe, you aren't concerned about the difference between 5.92 percent annual growth and 5.93 percent.

To store huge datasets effectively, we've seen a new breed of databases appear. These are frequently called NoSQL databases, or Non-Relational databases, though neither term is very useful. They group together fundamentally dissimilar products by telling you what they aren't. Many of these databases are the logical descendants of Google's BigTable and Amazon's Dynamo, and are designed to be distributed across many nodes, to provide "eventual consistency" but not absolute consistency, and to have very flexible schema. While there are two dozen or so products available (almost all of them open source), a few leaders have established themselves:

  • Cassandra: Developed at Facebook, in production use at Twitter, Rackspace, Reddit, and other large sites. Cassandra is designed for high performance, reliability, and automatic replication. It has a very flexible data model. A new startup, Riptano, provides commercial support.
  • HBase: Part of the Apache Hadoop project, and modelled on Google's BigTable. Suitable for extremely large databases (billions of rows, millions of columns), distributed across thousands of nodes. Along with Hadoop, commercial support is provided by Cloudera.

Storing data is only part of building a data platform, though. Data is only useful if you can do something with it, and enormous datasets present computational problems. Google popularized the MapReduce approach, which is basically a divide-and-conquer strategy for distributing an extremely large problem across an extremely large computing cluster. In the "map" stage, a programming task is divided into a number of identical subtasks, which are then distributed across many processors; the intermediate results are then combined by a single reduce task. In hindsight, MapReduce seems like an obvious solution to Google's biggest problem, creating large searches. It's easy to distribute a search across thousands of processors, and then combine the results into a single set of answers. What's less obvious is that MapReduce has proven to be widely applicable to many large data problems, ranging from search to machine learning.

The most popular open source implementation of MapReduce is the Hadoop project. Yahoo's claim that they had built the world's largest production Hadoop application, with 10,000 cores running Linux, brought it onto center stage. Many of the key Hadoop developers have found a home at Cloudera, which provides commercial support. Amazon's Elastic MapReduce makes it much easier to put Hadoop to work without investing in racks of Linux machines, by providing preconfigured Hadoop images for its EC2 clusters. You can allocate and de-allocate processors as needed, paying only for the time you use them.

Hadoop goes far beyond a simple MapReduce implementation (of which there are several); it's the key component of a data platform. It incorporates HDFS, a distributed filesystem designed for the performance and reliability requirements of huge datasets; the HBase database; Hive, which lets developers explore Hadoop datasets using SQL-like queries; a high-level dataflow language called Pig; and other components. If anything can be called a one-stop information platform, Hadoop is it.

Hadoop has been instrumental in enabling "agile" data analysis. In software development, "agile practices" are associated with faster product cycles, closer interaction between developers and consumers, and testing. Traditional data analysis has been hampered by extremely long turn-around times. If you start a calculation, it might not finish for hours, or even days. But Hadoop (and particularly Elastic MapReduce) make it easy to build clusters that can perform computations on long datasets quickly. Faster computations make it easier to test different assumptions, different datasets, and different algorithms. It's easer to consult with clients to figure out whether you're asking the right questions, and it's possible to pursue intriguing possibilities that you'd otherwise have to drop for lack of time.

Hadoop is essentially a batch system, but Hadoop Online Prototype (HOP) is an experimental project that enables stream processing. Hadoop processes data as it arrives, and delivers intermediate results in (near) real-time. Near real-time data analysis enables features like trending topics on sites like Twitter. These features only require soft real-time; reports on trending topics don't require millisecond accuracy. As with the number of followers on Twitter, a "trending topics" report only needs to be current to within five minutes -- or even an hour. According to Hilary Mason (@hmason), data scientist at bit.ly, it's possible to precompute much of the calculation, then use one of the experiments in real-time MapReduce to get presentable results.

Machine learning is another essential tool for the data scientist. We now expect web and mobile applications to incorporate recommendation engines, and building a recommendation engine is a quintessential artificial intelligence problem. You don't have to look at many modern web applications to see classification, error detection, image matching (behind Google Goggles and SnapTell) and even face detection -- an ill-advised mobile application lets you take someone's picture with a cell phone, and look up that person's identity using photos available online. Andrew Ng's Machine Learning course is one of the most popular courses in computer science at Stanford, with hundreds of students (this video is highly recommended).

There are many libraries available for machine learning: PyBrain in Python, Elefant, Weka in Java, and Mahout (coupled to Hadoop). Google has just announced their Prediction API, which exposes their machine learning algorithms for public use via a RESTful interface. For computer vision, the OpenCV library is a de-facto standard.

Mechanical Turk is also an important part of the toolbox. Machine learning almost always requires a "training set," or a significant body of known data with which to develop and tune the application. The Turk is an excellent way to develop training sets. Once you've collected your training data (perhaps a large collection of public photos from Twitter), you can have humans classify them inexpensively -- possibly sorting them into categories, possibly drawing circles around faces, cars, or whatever interests you. It's an excellent way to classify a few thousand data points at a cost of a few cents each. Even a relatively large job only costs a few hundred dollars.

While I haven't stressed traditional statistics, building statistical models plays an important role in any data analysis. According to Mike Driscoll (@dataspora), statistics is the "grammar of data science." It is crucial to "making data speak coherently." We've all heard the joke that eating pickles causes death, because everyone who dies has eaten pickles. That joke doesn't work if you understand what correlation means. More to the point, it's easy to notice that one advertisement for R in a Nutshell generated 2 percent more conversions than another. But it takes statistics to know whether this difference is significant, or just a random fluctuation. Data science isn't just about the existence of data, or making guesses about what that data might mean; it's about testing hypotheses and making sure that the conclusions you're drawing from the data are valid. Statistics plays a role in everything from traditional business intelligence (BI) to understanding how Google's ad auctions work. Statistics has become a basic skill. It isn't superseded by newer techniques from machine learning and other disciplines; it complements them.

While there are many commercial statistical packages, the open source R language -- and its comprehensive package library, CRAN -- is an essential tool. Although R is an odd and quirky language, particularly to someone with a background in computer science, it comes close to providing "one stop shopping" for most statistical work. It has excellent graphics facilities; CRAN includes parsers for many kinds of data; and newer extensions extend R into distributed computing. If there's a single tool that provides an end-to-end solution for statistics work, R is it.

Making data tell its story

A picture may or may not be worth a thousand words, but a picture is certainly worth a thousand numbers. The problem with most data analysis algorithms is that they generate a set of numbers. To understand what the numbers mean, the stories they are really telling, you need to generate a graph. Edward Tufte's Visual Display of Quantitative Information is the classic for data visualization, and a foundational text for anyone practicing data science. But that's not really what concerns us here. Visualization is crucial to each stage of the data scientist. According to Martin Wattenberg (@wattenberg, founder of Flowing Media), visualization is key to data conditioning: if you want to find out just how bad your data is, try plotting it. Visualization is also frequently the first step in analysis. Hilary Mason says that when she gets a new data set, she starts by making a dozen or more scatter plots, trying to get a sense of what might be interesting. Once you've gotten some hints at what the data might be saying, you can follow it up with more detailed analysis.

There are many packages for plotting and presenting data. GnuPlot is very effective; R incorporates a fairly comprehensive graphics package; Casey Reas' and Ben Fry's Processing is the state of the art, particularly if you need to create animations that show how things change over time. At IBM's Many Eyes, many of the visualizations are full-fledged interactive applications.

Nathan Yau's FlowingData blog is a great place to look for creative visualizations. One of my favorites is this animation of the growth of Walmart over time. And this is one place where "art" comes in: not just the aesthetics of the visualization itself, but how you understand it. Does it look like the spread of cancer throughout a body? Or the spread of a flu virus through a population? Making data tell its story isn't just a matter of presenting results; it involves making connections, then going back to other data sources to verify them. Does a successful retail chain spread like an epidemic, and if so, does that give us new insights into how economies work? That's not a question we could even have asked a few years ago. There was insufficient computing power, the data was all locked up in proprietary sources, and the tools for working with the data were insufficient. It's the kind of question we now ask routinely.

Data scientists

Data science requires skills ranging from traditional computer science to mathematics to art. Describing the data science group he put together at Facebook (possibly the first data science group at a consumer-oriented web property), Jeff Hammerbacher said:

... on any given day, a team member could author a multistage processing pipeline in Python, design a hypothesis test, perform a regression analysis over data samples with R, design and implement an algorithm for some data-intensive product or service in Hadoop, or communicate the results of our analyses to other members of the organization

Where do you find the people this versatile? According to DJ Patil, chief scientist at LinkedIn (@dpatil), the best data scientists tend to be "hard scientists," particularly physicists, rather than computer science majors. Physicists have a strong mathematical background, computing skills, and come from a discipline in which survival depends on getting the most from the data. They have to think about the big picture, the big problem. When you've just spent a lot of grant money generating data, you can't just throw the data out if it isn't as clean as you'd like. You have to make it tell its story. You need some creativity for when the story the data is telling isn't what you think it's telling.

Scientists also know how to break large problems up into smaller problems. Patil described the process of creating the group recommendation feature at LinkedIn. It would have been easy to turn this into a high-ceremony development project that would take thousands of hours of developer time, plus thousands of hours of computing time to do massive correlations across LinkedIn's membership. But the process worked quite differently: it started out with a relatively small, simple program that looked at members' profiles and made recommendations accordingly. Asking things like, did you go to Cornell? Then you might like to join the Cornell Alumni group. It then branched out incrementally. In addition to looking at profiles, LinkedIn's data scientists started looking at events that members attended. Then at books members had in their libraries. The result was a valuable data product that analyzed a huge database -- but it was never conceived as such. It started small, and added value iteratively. It was an agile, flexible process that built toward its goal incrementally, rather than tackling a huge mountain of data all at once.

This is the heart of what Patil calls "data jiujitsu" -- using smaller auxiliary problems to solve a large, difficult problem that appears intractable. CDDB is a great example of data jiujitsu: identifying music by analyzing an audio stream directly is a very difficult problem (though not unsolvable -- see midomi, for example). But the CDDB staff used data creatively to solve a much more tractable problem that gave them the same result. Computing a signature based on track lengths, and then looking up that signature in a database, is trivially simple.

Hiring trends for data science

It's not easy to get a handle on jobs in data science. However, data from O'Reilly Research shows a steady year-over-year increase in Hadoop and Cassandra job listings, which are good proxies for the "data science" market as a whole. This graph shows the increase in Cassandra jobs, and the companies listing Cassandra positions, over time.

Entrepreneurship is another piece of the puzzle. Patil's first flippant answer to "what kind of person are you looking for when you hire a data scientist?" was "someone you would start a company with." That's an important insight: we're entering the era of products that are built on data. We don't yet know what those products are, but we do know that the winners will be the people, and the companies, that find those products. Hilary Mason came to the same conclusion. Her job as scientist at bit.ly is really to investigate the data that bit.ly is generating, and find out how to build interesting products from it. No one in the nascent data industry is trying to build the 2012 Nissan Stanza or Office 2015; they're all trying to find new products. In addition to being physicists, mathematicians, programmers, and artists, they're entrepreneurs.

Data scientists combine entrepreneurship with patience, the willingness to build data products incrementally, the ability to explore, and the ability to iterate over a solution. They are inherently interdiscplinary. They can tackle all aspects of a problem, from initial data collection and data conditioning to drawing conclusions. They can think outside the box to come up with new ways to view the problem, or to work with very broadly defined problems: "here's a lot of data, what can you make from it?"

The future belongs to the companies who figure out how to collect and use data successfully. Google, Amazon, Facebook, and LinkedIn have all tapped into their datastreams and made that the core of their success. They were the vanguard, but newer companies like bit.ly are following their path. Whether it's mining your personal biology, building maps from the shared experience of millions of travellers, or studying the URLs that people pass to others, the next generation of successful businesses will be built around data. The part of Hal Varian's quote that nobody remembers says it all:

The ability to take data -- to be able to understand it, to process it, to extract value from it, to visualize it, to communicate it -- that's going to be a hugely important skill in the next decades.

Data is indeed the new Intel Inside.




O'Reilly publications related to data science

R in a Nutshell
A quick and practical reference to learn what is becoming the standard for developing statistical software.

Statistics in a Nutshell
An introduction and reference for anyone with no previous background in statistics.

Data Analysis with Open Source Tools
This book shows you how to think about data and the results you want to achieve with it.

Programming Collective Intelligence
Learn how to build web applications that mine the data created by people on the Internet.

Beautiful Data
Learn from the best data practitioners in the field about how wide-ranging -- and beautiful -- working with data can be.

Beautiful Visualization
This book demonstrates why visualizations are beautiful not only for their aesthetic design, but also for elegant layers of detail.

Head First Statistics
This book teaches statistics through puzzles, stories, visual aids, and real-world examples.

Head First Data Analysis
Learn how to collect your data, sort the distractions from the truth, and find meaningful patterns.




1 The NASA article denies this, but also says that in 1984, they decided that the low values (whch went back to the 70s) were "real." Whether humans or software decided to ignore anomalous data, it appears that data was ignored.

2 "Information Platforms as Dataspaces," by Jeff Hammerbacher (in Beautiful Data)

3 "Information Platforms as Dataspaces," by Jeff Hammerbacher (in Beautiful Data)



Sent from my iPad