2012-09-24

PUE, CO2 and NYT

Crater Lake Tour 2012

The NY Times has published an article "exposing" the shocking power wasted in a datacentre. It's an interesting read, even if their metrics "1.8 trillion gigabytes" take work to convert into meaningful values, which, assuming they use the HDD vendor's abuse of the values G and T in their disks specs, work out as:
"2,000 gigabytes": ~2TB
"50,000 gigabytes": ~50TB.
"roughly a million gigabytes": ~1 PB.
"1.8 trillion gigabytes": ~1.8 Exabytes.
"76 billion kilowatt hours: 76e9 KWh = 76e6 MWh = 76e3 GWh

There's already a scathing rebuttal, which doesn't say much I disagree with.

One part of the NYT article involved looking round a "datacenter" and discovering lots of unused machines, services that only get used intermittently. I'm assuming this is some kind of enterprise datacentre, a room or two set up a decade ago to host machines. Those underused machines should be eliminated; their disk image converted to a VM and then hosted under a hypervisor. Result: less floor space, CPU power and HDD momentum wasted.

Those enterprise datacentres are the ones whose PUE tends to be pretty bad -because it's mixed in with the rest of the site's aircon budget, and not as significant & visible a cost as it is for the big facilities. Google, Amazon and Facebook do care about this; they are probably the people backing the ARM-based servers, such as those running Hadoop jenkins builds. What those vendors care about tends to be cost though: cost of HW, cost of power, cost of land, cost of packets.

What the article doesn't look at -but the folks at MastodonC will presumably cover at Strata EU- is not the energy cost of computation, but the CO2 cost, Those datacentres in VA, where Amazon US-East is, have awful CO2 footprints, being all coal-powered. That's why it's ironic that the NYT complains about Amazon's diesel generators being pollution -in a part of the world where mountain-top mining converts entire mountains into smoke. They'd have been better off looking at the CO2 footprint of the datacentres, and of the other industries in the area.

MastodonC's dashboard is why I'm storing data and spinning up t1.micro instances in US-West 2 -Oregon; lowest CO2 footprint of their US sites.

I was also kind of miffed as the paper's criticism of power lines "financed by millions of ordinary ratepayers". Surely freeways were "financed by millions of ordinary ratepayers", yet the NYT has never done a shocking critique of Walmart's use of them to ship consumer goods round in fuel-inefficient diesel trucks, despite the fact an energy efficient alternative (electric trains) have existed for decades.

One thing the NYT does hint at is the storage cost -and hence the power cost- of old email attachments. It makes me think that I should clean some of the old junk up. What they don't pick up on is the dark secret of Youtube: the percentage of videos that are of cats. If you want someone to blame, blame the phones that make taking such videos trivial, and the people who upload them.

[Photo: Crater Lake, OR. The sky is hazy as the forest fires in Lassen and west of Redding are bringing smoke up from CA].

2012-09-17

My Hadoop-related Speaking Schedule

I'm back from the US, where I had lots of fun getting the HA HDP-1 stuff out the door -I know about Linux Resource Agents, and too much about Bash -though that knowledge turns out to be terrifyingly useful.

Here's a pic of me sitting outside a cabin in Yosemite Valley where we spent a couple of nights -Camp 4 wasn't on the permitted accommodation list this time.
Curry Camp Cabin, Yosemite

Some people may be thinking "cabin?" "Yosemite?" and "Isn't that where all those people caught Hantavirus and died?". The answer is yes -though they were in wooden-walled tent-things about 100 metres away, and the epidemology assessments show that even for them the risk is very small. The press like headlines like "20,000 people may be at risk" -missing the point that the larger the set of people "present" for the same number of "ill", the smaller P(ill | present). Which is good as P(die | ill)=0.4.

Even so,  I've had some good discussions with the family doctor and the UK Health Protection Agency, who did write a letter saying "if you show symptoms of flu within 6 weeks of visiting, get to a hospital for a blood test". As the doctor said "we don't get many cases of Hantavirus in Bristol", so it's not something they are geared up for. You know that when they start looking at the same web pages you've already read.

Well, we've got 1-2 weeks left to go. And it was excellent in Yosemite, though next time I'd stay more in Tolumne Meadows than in the valley itself (too busy), and maybe sort out the paperwork to go back-country. 

Yosemite

Assuming that I remain alive for the next fortnight, here are where I'm going to be speaking over the next few months.

Strata EU: Data Availability and Integrity in Apache Hadoop.

I've already done a preview of the talk at a little workshop in Bristol -the live demo of RHEL HA failover did work, so I hope to repeat it. I'll be manning the Hortonworks Booth and wearing branded T-shirts, so will be findable -though I plan to attend some of the talks. In particular, one of the people behind Spatial Analyis UK will be talking -and I just love their maps.

Big Data Con London, Hadoop as a Data Refinery.

Here I'll be exploring the "Data Refinery" metaphor as a way to visualise and communicate the role of the Hadoop stack in existing organisations.

ApacheCon EU, Introduction to Hadoop-dev.

I'm going to talk about the Hadoop development process, QA and testing, contributions. This isn't going be a basic "here's SVN", or a "Hortonworks and Cloudera can handle everything" talk, but one that looks at the current process -both strengths and weaknesses. As a committer who was not only on their own for some years, but still in a different TZ, I know the problems that arise. I believe it is essential for people using Hadoop in the field to get their feedback in, through JIRA, tests & patches. If there is one thing that I think needs work is to have a semi-formalised process for external projects to do mentored work relating to Hadoop. That's companies, individuals, interns and university research. All to often we don't know that someone is working on a feature until they turn up with something big that cuts across the projects -and at that point it's too late to shape, to open up to external input, or to even comprehend. Just as apache has an incubator, I think we need something structured -as the alternative is that this work falls on the floor and ends up wasted.


2012-08-30

RE: Hello from Twitter!

My life is less interesting than it seems on facebook

Yet another LI approach, this one that didn't know who at Twitter was a Hadoop committer and therefore someone I knew/had consumed beer at the Highbury with.

What I will flag up here is Twitter's storage team's plans to build a new data platform from scratch. Not heard of that before.


Hi. I 'm afraid I must decline your invitation. I am having lots of fun at hortonworks building the future of Hadoop, and am not interested in any discussions.

While I am at it, can I point out
http://steveloughran.blogspot.co.uk/2012/07/hadoop-recruiter-minefield.html
and
http://steveloughran.blogspot.co.uk/2012/07/reminder-to-hadoop-recruiters-do-your.html

these state my reference policy for dealing with unsolicited approaches. A quick scan of the Hadoop committer list would identify who I have worked with at twitter, so who to approach me through more personally.

On 8/30/12 1:00 AM, Mike  wrote:
--------------------
Hi Steve,

I lead passive candidate engagement for the core storage team at Twitter. I ran across your profile and I wanted to see if you may be open to a short conversation.  The storage team is working on a next generation big data platform we are building from scratch.  This will fundamentally change the way we store and analyze data.  Would you be available to speak sometime this week?

I look forward to your response.

Cheers,

Mike

2012-08-24

EPO & TdF


Summit Approach
I guess I should undeclare all photos of me cycling in either US postal or T Mobil cycling tops, such as here where A. and I topped out Col de Madeleine (1993m) in 2009, having a meal at the top before descending -the first time my son had been over 40mph on a bicycle before. Those burley tagalongs corner well -the rack mount gives them a good CofG, but are not so good off road.

Regarding the TdF, I was down in Grenoble in '98, and almost drove up to Annecy to catch the stage there -I chose to head east and ride Galibier from both sides instead. That was a good choice, as that was the day the riders had a sit down strike "If you make the drug tests harder -make the tour easier".

This shows the problem. The riders want to win -but the TV companies wanted exciting television; the little towns wanted the intermediate sprints, as did that French betting company (PMC?), the sponsors wanted the TV coverage. EPO and blood doping were the dirty little secrets: undetectable drugs that meant from Indurain's era onwards, all winners of the TdF probably cheated.

It's hard to blame or criticise Lance Armstrong here, because, well, that was how the game went that year -and it wasn't just the cyclists who benefited.

Summit

Me: I just regret not taking an afternoon out from Geneva in 1988 to catch Lemond.

2012-08-15

Page Mill Hill Work, T+20

This is me atop Page Mill; photo on Ilford B&W, hand developed:
Pagemill Summit 1992

Date? Summer 1992; spending a few weeks in Santa Clara.

I've just made it up Page Mill, the bike -my original mountain bike- is carrying about 6-8kg of surplus weight on the back.


Here's the scene again, this time on a digital camera that even knows where it is:
20 years on

Date: Summer 2012; spending a few weeks in Sunnyvale.


The 6-8kg of surplus luggage has moved off the bike into my body, making it impossible for me to leave it behind on rest days; the fringe on my hair has moved back, and the sunglasses hide the fact I look more tired.

Otherwise, not much has visibly changed.

Except that climb up Page Mill on the first photo was followed by a ride all the way up Skyline, dropping down to the pacific coast side of San Francisco where I eventually turned up at my motel to discover that my reservation had been voided & because of some event in the city if I wanted somewhere to stay I would have to sprint over to the YMCA in that fairly rough part off Market -Tenderloin?- where before I could get a room the people in front were complaining they'd been robbed through a window on the third floor. Needless to say the bike came in the room. Estimated distance: 90-100 miles, 3000-5000' of ascent.

The next day: over into Marin, over Mt Tamalpais on an off-road trail, where I got to enjoy overtaking people who didn't have luggage, then over to Pt Reyes and the Youth Hostel there; 70-80 miles, 5000-6000' of ascent.

These days, up Pagemill then S. on Skyline before descending to Cupertino and home is enough for me to declare victory (50 miles; 3000' of up), where I can then settle down and drink a beer pretending to myself that I've earned it.

That's the difference.

2012-07-31

Welcome to Chaos

Welcome to Bristol

Netflix have published their "Chaos Monkey" code on Github; ASL Licensed. I have already filed my first issue, having looked through the code -an issue that is already marked as fixed.

Netflix bring to the world the original Chaos Monkey -tested against production services.

Those of us playing with failures, reliability and availability in the Hadoop world also need something that can generate failures, though for testing the needs are slightly different:
  1. Failures triggered somewhat repeatedly.
  2. Be more aggressive.
  3. Support more back ends than Amazon -desktop, physical, private IaaS infrastructures.
#1 and #2 are config tuning -faster killing, seeded execution.

#3? Needs more back ends. The nice thing here is that there's very little you need to implement when all you are doing is talking to an Infrastructure Service to kill machines; the CloudClient interface has one method:
  void terminateInstance(String instanceId);
That needs to be aided with something to produce a list of instances, and of course there's the per-infrastructure configuration of URLs and authentication.


My colleague Enis has been doing something for this for HBase testing.; independently I've done something in groovy for my availability work, draft package, org.apache.chaos. I've done three back ends :
  1. SSH in to a machine and kill a process by pid file.
  2. Pop up a dialog telling the user to kill a machine (not so daft, good for semi-automated testing).
  3. Issue virtualbox commands to kill a VM.
All of these are fairly straightforward to migrate to the Chaos Monkey; they are all driven by config files enumerating the list of target machines, plus some back-end specific options (e.g. pid file locations, list of vbox UUIDs).

Then there's the other possibilities: VMWare, fencing devices on the LAN, ssh in and issue "if up/down" commands (though note that some infrastructures, such as vSphere, recognise that explicit option and take things off HA monitoring). All relatively straightforward. 

Which means: we can use the Chaos Monkey as a foundation for testing how distributed systems, especially the Hadoop stack components, react to machine failover -across a broad set of virtual and physical infrastructures. 


That I see the appeal of.


Because everyone needs a Chaos Monkey.

[update 13:24 PST, fixed first name, thank you Vanessa Alvarez!]

2012-07-26

Reminder to #Hadoop recruiters: do your research

It was only last week that I blogged about failing Hadoop recruiter approaches.

Key take-aways were
  • Do your research from the publicly available data.
  • Use the graph, don't abuse it.
  • Never try to phone me.
Given it is fresh in my blog, and that that blog is associated with my name and LI profile, it's disappointing to see that people aren't reading it.

Danger Vengeful God

A couple of days ago, someone trying to get me on Twitter looking for SAP testers, which is as relevant and appealing to me as a career opportunity in marketing.

Today, some LI email from Skype.com
From: James
Date: 26 July 2012 07:35
Subject: It's time for Skype.
To: Steve Loughran

Dear Steve,

Apologies for the unsolicited nature of the message; it seemed the most confidential way to approach you, although I shall try to contact you by phone as well.

I am currently working within Skype's Talent Acquisition team and the key focus in my role is to search and hire top talent in the market place. I am currently looking at succession planning both for now but also on a longer term plan. I would be keen to have a conversation with you about potential opportunities, and introduce myself and tell you more about Skype (Microsoft) and our hiring in London around "Big Data".

I look forward to hearing from you.
James
See that? He's promising to try and contact me by phone as if I'm going to be grateful. Yet last week I stated that as my "do that and I will never speak to you" policy.  Nor do I consider unsolicited emails confidential -as you can see.

I despair.

The data is there, use it. If you don't, well, what kind of data mining company are you?

For anyone who does want an exciting and challenging job in the Hadoop world, one where your contributions will go back into open source and you will become well known and widely valued, can I instead recommend one of the open Hortonworks positions? We. Are. Having. Fun.

As an example, here is a scene at yesterday's first birthday BBQ:
Garden Party


Eric14, co-founder and CTO on the left, Bikas on the right, wearing buzz lightyear ballons after food and beverages; nearby Owen is wearing facepaint. Bikas is working on Hadoop on Windows, and has two large monitors showing Windows displays in his office, alongside the Mac laptop -the outcome of that work is not just that Hadoop will be a first class citizen in the Windows platform, you'll get excellent desktop and Excel integration. Joining that team and you get to play with this stuff early -and bring it to the world.
 
I promise I will not phone anyone up about these jobs.