2012-05-19

Possible Stonebraker Trajectories

City Farm, St Werburgh's

In my time I've managed to say bad things about SQL in an email discussion that had the (now officially late) Jim Gray in the CC:. I also faulted the first generation of grid engines for believing the storage vendors when they said "don't worry about storage" when I was participating in a panel with the lead of the Condor project. Despite this, the ACM doesn't give me a blog page.

It's a shame, then, to see the ACM letting Stonebraker publish another of his rants, Possible Hadoop Trajectories

First he gets a dig in at Java developers "discovering parallel processing". Actually, they've had threads and small clusters for a long time actually. What Hadoop brings to the planet is the opening up of the ability to work with thousands of cores and PB of data to lots of people. The query languages, Pig and Hive, mean that you don't need to learn Java either -any more than you need Objective C skills to use an iPhone.

Then he tries the "it's so inefficient the planet dies" story. At least they are planning to build their next supercomputer by a hydro plant -but it's still going to have a power budget of Megawatts, so calling Hadoop environmentally unfriendly is ironic. It's like Airbus saying their planes are greener than Boeings' -when both are flying people across the Atlantic(*).

If there's one thing that really annoys me here it's a bit from the opening paragraph:
"we applaud Hadoop for its success in this area, which we believe is due largely to the simplicity and accessibility of its environment. "
Exactly.
  1. It is simple to use because Map(x):(k,y) and Reduce(k, [y1,...,yn]): z are easy to understand and play with. You can write efficient routines without knowing relational algebra and set theory, unlike, say RDBMs [Codd71].
  2. It is accessible because it runs (slowly) on your laptop and massively faster on your production cluster. It is also economically accessible because it is free to download and start to play with. You may need training or support -which Hortonworks will gladly offer -but that payment is optional. You can learn through the books and trial and error. You can learn to support your own system. Doing so does give you the duty to rummage through the code yourself, but if you contribute any fixes you have made back, even your in-house support efforts benefit the community as a whole.
There is nothing wrong with simplicity and accessibility. This is why PHP is one of the key development platforms of Facebook. When Facebook wanted those PHP developers to work with Hadoop, they didn't say "go learn Java". They said "here's an SQL bridge", called Hive. For those people who already known SQL, Hive lets you work with Hadoop without having to write a line of Java. There is nothing wrong with that and it does not make sense to denounce Hadoop because someone wrote tooling to help SQL experts work with it.

That does not mean that SQL is a good language. That little fact has been forgotten since RDMBs's became widespread, when developers learned to write things like "SELECT * FROM users WHERE name="steve"". SQL is a language designed to make script injection the default operation; something SELECT * FROM users where name=""; DROP TABLE users".

SQL started out as SEQUEL: "Structured English Query Language". It was written on the expectation that business people -presumably the same people that COBOL was targeted at- would sit at their shiny new IBM teletype and type in an 'ad-hoc query'. That's right: SQL was not targeted at developers, but "normal people" -and to be easy for COBOL and PL/I developers to embed. A key goal of the SQL language was to present the same capabilities, and a consistent syntax, to users of the PL/I and COBOL host languages and to ad hoc query users.[Chamberlin81].

Nowadays, the main experts in SQL are people like Facebook's PHP devs, and script hackers. Java developers run from it, hence the broad set of O/R mapping tools. Enterprise Java Beans were first; someone had a vision that people would write reusable "beans" to represent enterprise entities (User, Customer, Purchase), and that there would be some kind of market for that. Well, that died, but Hibernate and Spring keep letting Java devs write distributed database transactions without having to learn SQL. Where are Stonebraker's snide language-elist comments then? Why no ACM article saying "ORM tools have finally brought the power of the database to Java developers"? Is it because he felt that ORM was a good idea, or that he recognizes that tools to make working databases easier benefited him?

The harsh truth is that SQL is not a particularly good language for expressing relations and predicates. Back in 1984 the illustrious C.J Date (as in "Introduction to Database Systems" Date) published a 47 page dot-matrix-printed critique of the language [Date84] -an article whose criticism on the difficulty embedding SQL into PL/I is effectively the precursor to all critiques of O/R mapping. It's SQL/Language mapping, and there've been problems mixing code and SQL back since System R first booted. A key problem is that all it does is read and write data from the DB, but for programs you need more than that, so you end up mixing SQL queries in that COBOL-esque syntax with the real code, either through some contrived ORM process or some hand-rolled string construction thing that at best is a maintenance task and at worst leaves your entire site's credit card records up on pastebin.

If you did want to work with databases properly, you'd need a programming language which makes relations and predicate calculus integral parts of the language: Prolog, Linq and, effectively, Erlang. Linq interests me as it is the most recent attempt, and because Dryad/Linq showed that it could do more than just database lookups.

Returning to System R, the database from which DB2 and Oracle DB are derived, [Chamberlin81] concludes with a lovely sentence:
We feel that our experience with System R has clearly demonstrated the feasibility of applying a relational database system to a real production environment.
Which can be translated as: "even though people preferred more efficient low-level data storage techniques, hand-tuned for the specific application, pre-written in assembly language, COBOL or PL/I, the System R team -including the illustrious and now sadly absent Jim Gray- felt that making working with data easier outweighed the alternative.

That's something Stonebraker appears to have missed. The RDBMs isn't an end in itself -it's a means to an end. A tool. As is Hadoop. A tool to let you work with data at a scale and price point that that the commercial RDBMs can't play at.

Is MapReduce the meta-algorithm to solve everything? Of course not. The Stratosphere team in Berlin, the Asterix team at UC Berkeley are key leaders in the academic space -there a both ideas and code to pick up here. Then there's the real world projects coming out of the web companies, who do have to work at a scale and price point that RDMBs's can't match: Pig, HBase, Hama, Giraph, S3; other key-value stores nearby: Cassandra, Project Voldemort. All of these worked for their organizations.

Which is why I have a quote; a slight mutation of the system R conclusion based on the experiences of all the Hadoop users:
We feel that our experience with Hadoop has clearly demonstrated the feasibility of applying a Hadoop system to a real production environment.
For anyone interested in things like Stratosphere, the Graph Layer, what Yarn allows &c, there's a two day workshop after Berlin Buzzwords, "Beyond MapReduce" -free for all conference attendees. Stonebraker is cordially invited to attend the conference and the workshop. I'll gladly sit next to him on a panel and say things he won't agree with.

(*) This post was written on an A340-400 between SFO and LHR. I do have all the cited papers on my laptop. If you are going to argue with the RDBMs people, you need to know where they are coming from.

[Chamberlin81] D Chamberlin et al., A History and Evaluation of System R, 1981.
[Codd71] E. F Codd, A Database Sublanguage Founded on the Relational Calculus, 1971
[Date84]: C.J. Date, A Critique of the SQL Database Language 1984

2012-04-19

Pythonium, or why is there no javash?


Palmer Ray Solicitors

One thing I've been doing this month is learning some basic Python. Not because I have vast urge to switch all my code to a typeless language just to apply the lambda calculus to arrays (though that appeals, obviously).

No, the reason I want to learn Python can be summarised in one word: bash.

All Java projects I've touched have tended to have some startup bash script, maybe also a windows .BAT file if really needed. And those bash scripts suck.

They are written by Java developers who don't know/understand bash, and contain lots of invalid assumptions and bad code. I know, I fit into that category.

They are invariably brittle towards: spaces in paths and filenames, things not being where they should be, environment variables not being set.

They are hard to port across platforms because bash relies on lots of unix programs to do a lot of work, programs that take different arguments on Linux, MacOS and legacy SunOS platforms. (ooh, SunOS was built on BSD which is based on APIs from AT&T. I hope nobody involved in lawsuit about API copyright notices that).

They don't test within the Java JUnit framework world. You can do it, but it's hard and doesn't stress the troublespots -env variables, spaces in paths, different programs and args. This means that the various if [ ] ; queries to handle platform specificness don't get a look in before release.

As a result, those little startup scripts are an inordinate amount of trouble -probably the most brittle and lowest quality bit of source within a Java project. (second most if there are .BAT/.CMD scripts too).

Yet every Java project ends up having one. Why? There's no easy way to launch a java process with the classpath set up right, with parameters passed down and actually working.

Yes, you could double click on a JAR file with a (relative) classpath in the manifest, but how does that pass down -Xmx16g -server -XX:+TieredCompilation-XX:+UseCompressedOOps -XX:UseNUMA -XX:+UseParallelGC XX:+UseParallelOldGC -Dlog4j.properties=/var/app2/conf/log4j.properties? It doesn't. Hence: the shell script.

In an ideal world. you wouldn't need it. You could have some .jar-opts file that would live alongside the JAR file that would have all these options, one to a line. You could have that .jnlp stuff actually do something useful except trying to load JavaFX on demand for all three people that are trying to run a JavaFX-based applet.

Or, and this would be nice, you could have a shell runtime that compiled and executed .java files. Mark a file as executable, put a #!/bin/java on the first line, and it could be compiled and executed on demand. Then you could have java startup code written in Java, setting up the main program.

Adding a javac compiler wouldn't add to the runtime footprint much. There's already javafx, javascript, the whole of the JRE library packages. A basic compiler is not that much space.

On devices with limited resources -storage, cpu- you could argue against on-demand compilation, but for everything else, just compile it as needed. While people doing commercial code may worry about this, if you are doing modern webapps, your source is being downloaded to every browser in the .js files, and the server-side stuff is hidden. Oh, and decompilers show us stuff anyway.

In the OSS world, having the source on everyone's machine is a tangible benefit -it gets the source into the hands of the users, no way of hiding it from the people who may have to take up the maintenance task in the future.

but no, there is no javash. Which is why I'm learning Python: so I have a language that isn't Bash or Perl to do the stuff that Java won't let me as it retains the old "compile then distribute" world view.

I don't need to learn much, just enough to write entry points that exec things or tell users off. And with ipython and the Python support in IntelliJ IDEA, that's fairly straightforward. It's certainly fun not having to wait for that compiler delay. Which is another issue: pre-compilation may save end-user time,  but it wastes developer time. As  a developer, I know what I value more, at least during the dev-and-test cycle.

[Photo: Stokes Croft storefronts]

2012-04-05

Joining Hortonworks to evolve #hadoop

Mountain Biking the Black Mountains

I've left HP. I did that on Monday, enjoying a final beer at lunchtime with my soon-to-be-ex colleagues, then heading home for a few weeks of parental responsibilities during the easter break.

Later this month, I will start work at Hortonworks, pushing the Hadoop stack forwards. I am really excited about this -I know a lot of people in the company already, and it's going to be great working with them!

Although the phrase "Big Data" is getting overused, it's obvious to me that there is a real coming together of different trends to make the whole Hadoop-based ecosystem as transformational as web servers were.
  1. There are so many devices in the modern world acting as data sources -physical devices such as mobile phones and jet engines, services such as web applications, people making use of devices and services. 
  2. In the past less data was generated -and it was normally thrown away. Too expensive to store, no perceived value.
  3. The cost per TB of HDD has fallen such that you can now afford to keep that data for later analysis
  4. You can't analyse it on single servers as the bandwidth of HDDs hasn't increased at the same rate as the storage capacity.
  5. The performance of a single CPU has effectively topped out too. All that is coming is more cores, more operations/joule (hopefully), different forms of parallel computation. The free speedups that the CPU vendors used to dish out are over. It's either single-machine parallelism or multi-machine. Oh, and either way: heterogeneity of some form or other.
  6. That means everyone is going to have to embrace parallel computing, on the single machine or in the rack -and with the right algorithms,  that rack can be made to deliver linear and sometimes superlinear speedup.
  7. If you want to work with the big datasets that you can collect today, you are going to need a rack of servers and a framework to let you process the data.
  8. The Hadoop platform provides the framework to store the data across those hard disks, and to distribute the work across them. It is becoming the single open-source alternative to Google's internal platform.

Where the future gets really interesting is that the Hadoop ecosystem provides those core services of a distributed computing platform: bulk storage (HDFS), scheduling (MRv2),  distributed state (ZooKeeper), integration with existing infrastructure (flume, squoop, Hive). These services can be used to build applications in and above Hadoop -HBase and Giraph are key examples; Cassandra a welcome friend. Big Data is the immediate reason to move into world, but ultimately it's Big Datacentre -not things like Java EE7 that just seem, well, so very last-century.

That's why I'm joining Hortonworks -to go full time on building the future platform for server-side computing.

[photo: preparing to descend into Crickhowell, Wales, 2011]

2012-03-16

Safety in High Energy Physics and US Visitor paperwork

The NYT article on US immigration and ESTA makes me think that I should publish the safety lecture from the November 2010 Bristol Hadoop workshop, which was hosted by Bristol University Physics Dept. All signs in the slides are from their physics building, except for the last slide.




The last form is genuine it's hosted on the state department as the 1405-0134 form, I've had to fill it in a couple of times. Things I like about it
  • I can cross more countries on a one day alpine bike ride than there is room for in the "countries you have visited" section. 
  • Giving to a charity is clearly something they don't expect people to do much of -again, space for two or three, and no time limit on how far back you must list your donations.
  • In places like Boise, Idaho, having firearms training is something they ought to give visitors, not ask if they have it.
Because form explicitly says "nuclear experience", the HEP folk can get away with saying no. Saying you work with antimatter or neutron beams is not the thing you want to do at an immigration border.

While taking the on-site photos I got cornered by the University site security people for taking pictures with an SLR as it is "what the police warned them about". I didn't point out to them that if I wanted to take photos discreetly I'd use a camera phone or my HD-resolution cycle helmet cam on the bike helmet I'd carry nonchalantly under one arm -as that would only make them think I was planning something. Better to stick to the idea that enemies of the state use SLRs -so anyone with an SLR is potentially an enemy. Anyway, I didn't argue, just sat there, let them look at the photos while verifying that I was visiting the physics dept. At some point the chief minion started talking about deleting one of the paper sign listing how many gigabequerels they had. Ignoring the fact that such info is available online, having that photo deleted would have wasted 15 minutes of my life once I got back home and undeleted it. At least they didn't ask to look at the laptop, as that would have created conflict.

Returning to ESTA, here's something funny about it. There is no online way to see if it expires. Apparently you get a renewal email, but I've never seen one, and last november I tried to see if mine was still valid before flew to the US. I didn't get a reply until after I'd flown out:


Dear Stephen,
I am sorry we were not able to respond to your question sooner. Hopefully you did not have any problems traveling to the US, but please write back if you still need help.

That is -we hope it hadn't expired yet because you'd have been stuffed if it had.

2012-03-14

Hadoop in Cloud Infrastructures

Ranier descent

People say "should you run Hadoop in the cloud?". I say "it depends".

I think there is value in Hadoop-in-cloud, I talked about doing it in 2010 at Berlin Buzzwords 2010; since then I've had more experience with using Hadoop and implementing cloud infrastructures.
  1. If your data is stored in a cloud provider's storage infrastructure, doing the analysis locally is the only rational action. It's that "work near the data" philosophy.
  2. If you are only doing some computation -say nightly- then you can rent some cluster time. Even if compute performance is worse, you can just rent some more machines to compensate.
  3. You may be able to achieve better security through isolation of clusters (depends on your IaaS vendor's abilities).
  4. No upfront capex; fund from ongoing revenue.
  5. Easier to expand your cluster; no need to buy more racks, find more rack space.
  6. You don't need to care about the problems of networking.
  7. Less of a problem of heterogenous clusters if you expand later.

Against that

  1. Cost of data storage can only increase at a rate proportional to ingress/retention rates.
  2. Cost of cluster time increases at a rate proportional to analysis performed. There is no "spare cluster time" for low priority work.
  3. Even if CPU time can scale up, IO rate of persistent data may not.
  4. Hadoop contains lots of assumptions about running in a static infrastructure; it's scheduling and recovery algorithms assume this.
Some examples of where Hadoop's assumptions diverge from that of cloud infrastructures:

  • HDFS assumes failures are independent, and places data accordingly (Google's Availability in Globally Distributed Storage Systems paper shows this doesn't hold in physical infrastructures, my notion of failure topologies expands on that)
  • MR blacklists failing machines, rather than releasing them and requesting new ones.
  • Worker nodes handle failure of master nodes by spinning on the hostname, not querying (dynamic) configuration data for new hostnames. Some of the HA HDFS may address that, I'm not tracking it enough.
  • Topology scripts are static. I've been slowly tweaking topology logic in 0.23+ but haven't put the dynamicness in there yet (HDFS and MR cache (name->rack) mappings on the assumption that the data is coming from slow to exec scripts, not fast & refreshable in-VM data).
  • Schedulers assume #of machines are static, don't allocate and release compute nodes based on demand and with knowledge of cost and quantum of CPU rentals. (I'm not sure quantum is the right term there, I mean the fact that VMs may be rented by the hour, 15 minutes, etc, so your scheduler should retain them for 59 minutes after acquiring them.
  • Scheduling doesn't bill different users for their cluster use in a way that is easily mapped to cluster time.

A lot of these are tractable, you just have to put in the effort. The Stratosphere team in Berlin are doing lots of excellent work here, including taking a higher level query language and generating an execution plan that is IaaS aware -you can optimise for fast (many machines) or lower cost (use less machines more efficiently).

In comparison, a physical cluster:
  • Offers a lower cost/TB of any corporate filestore to date other than people's desktop computers (which have a high TCO and power cost that is generally ignored), so enables you to store lots of stuff you would otherwise discard.
  • Let's you choose the hardware optimised for your current and predicted workloads.
  • Has free time for the low priority background work as well as the quicker queries that near-real-time UIs like.
  • May be directly accessible from desktops in the organisation (depends on security model of cluster).
  • Is easily hooked up to Jenkins infrastructure for execution of work as CI jobs.
  • Let's you do fancy tricks like striping of different MR versions across the racks for in-rack locality and different sets of task trackers for foreground vs background work, and different JTs (reduces memory use, cost of failure, etc).
  • Is way, way easier to hook up to internal databases, log feeds. To do ETL into your corporate oracle servers, you will need to run something behind the firewall to fetch it off the IaaS storage layer, rather than have your reducers push it to the RDBMS itself.
If you are generating data in house, in house clusters make a lot of sense.
This is why I say "it depends" -it depends on where you collect your data and what you plan to do with it.

As for the way Hadoop doesn't currently work so well in such infrastructures, well the code is there for people to fix. It's also a lot easier to test in-cloud behaviour, including resilience to failure, than it is with physical clusters.

[Photo: Descending Mt Ranier, 2000]

2012-02-17

The datacentre is the new laptop


Sepr on CK1 at PRSC

HP has just announced its forthcoming Gen 8 servers. Rather than go on about the usual stuff: CPUs, I/O bandwidth etc., or even the trend to put solid state storage off the PCIx bus, what's interesting to me is this: the servers are explicitly designed to be part of a larger system, a datacentre.

The existing products, they are individual servers you just happen to put into racks, you just happen to hook up to a switch you've stuck at the top of the same rack. The racks may be set up into hot rack/cold rack, but that's mostly a deployment detail the servers don't care about, except in ensuring airflow is good.

This has now changed.

The Datacenter as a computer argued that the software developers need to recognise that a datacentre is the new execution platform, one with mixed availability, limited bandwidth and other concerns that could be ignored before -or at least treated as the special case of "distributed systems", rather than what we have now: "systems". Everything is distributed.

These hardware changes mirror that. Here are some of the new concerns for both the ops team -and the applications themselves.
  1. Re-integration of Storage and Computation.
  2. Availability though replication. Less RAID-style hardware, more
    replication across machines.
  3. Inventory tracking -especially for identifying failure points, such as monitoring the history of specific batches of disks. If some appear particularly unreliable, you want to find all of them.
  4. Networking: 10 GbE is still a luxury, bonded 2x1 GbE good for availability too. Understanding network failures in data centers shows why ToR switches become the dominant network failure point in a cluster -and from a re-replication perspective, that's not ideal.
  5. Power management. Beyond just PUE, the metric of datacentre overhead, power consumption in the servers is a big concern.

The new servers then, are designed to live in this world.

Inventory They work out from the rack (don't ask me how, I don't know these things) where they are on it -information that can be propagated to the management tools so that they can be used for inventory tracking.

Networking Lots of ethernet ports. Some slow and inexpensive for management, faster ones for the application.

Power. This work here is something you can point to Chrandrakant Patel in HP Labs for. If you look at his published work, you can see a lot of it is about airflow and cooling in a datacentre. If you can improve that -as the container hosted datacentre pods can do- then your PUE is better. Why instrument the inside of the servers? It ensures that you can keep the hardware within its limits, because you have a better idea of what is going on inside. Every extra degree F, C or K you can take the air up, lets you save a lot of money over time. Yet the risk of overheating -and the cost of doing so- makes this dangerous. Knowing what is happening inside the servers give you more confidence of what's happening.

This is what the new servers enable. Which means that we are going from servers that you stack to servers that are designed to locate themselves in the racks, ideally hosted within a datacentre container that is optimised for airflow and designed to work as close to the limits of temperature as is considered safe based on the information coming out of the servers themselves.

Which is very close to what a laptop does: a box with optimised airflow and fans that come on when they feel it is important, and with a power budget that the system is designed to optimise. The datacentre is the new laptop, at least from a power and cooling perspective.


Now, what about the software? If the datacentre-level application infrastructure can get at the power, topology and network information, it could adapt itself better.

The topology information that the servers can determine could be used to dynamically generate the topology map for the cluster. It is entirely co-incidental that I'm typing this while my new topology patches are being tested in a adjacent console, but those changes (better support for topology sources other than the script runner, ability to dump the current topology) are effectively a precursor. I wouldn't do some fancy integrated java module though -better to have a topology source that just reads a java properties file and by polling for changes, can react to moving topologies. Let the management tooling generate that and it would propagate into HDFS and the RM/MR layer.

Power? If overheating is a problem, that server can be clocked back, which makes it slower. It may be better to actually tell the resource manager that there are less slots on that box, so reducing its actual workload. This could ensure that the work running in the remaining slots doesn't take longer than normal to complete.

Networking? We really need a way to get more information about the network backplane into the application -including the amount of bandwidth currently allocated to applications. Bandwidth can be a precious resource, but right now there is better tooling to manage it in a bittorrent client than there is between applications in a datacentre.

This is a challenge and an opportunity. A challenge: this information needs to be extracted and forwarded to the applications -which then need to act on it. An opportunity -it will make the applications and datacentres work better. Wave goodbye to writing topology scripts that don't work, say hello to being able to move servers around and have them the application infrastructure work out where they are. Worry less about uncontrolled backbone bandwidth use in a shared datacentre; have some policy tooling to manage it across applications. As for power, hope to see the electricity bills decrease.


[Artwork: Sepr on Jamaica Street, Stokes Croft]

2012-02-05

Just because you can rewrite your codebase doesn't mean you should have to



3DOM on Richmond Road: Remember the future

The strength of automated driven test suites is that you can verify that all the testable/covered parts of your system still work, even after major changes.

The strength of modern IDEs is at a click of mouse you can find wherever a class or method is used, so go to those places and edit them.

Even so: availability should not imply necessity: life is best when you don't have to use these features unless you really, really want to. Everything should work.

And usually it does. But not last week. A small bug surfaced on tuesday: you got an error marshalling strings with square brackets around, "[]". At first I did the obvious tactic: deny that this could be happening, but after repeated evidence to the contrary I sat down and had a look.

It turns out that the library, json-lib, with the same interface as the java Map interfaces, has an extra "feature" in that whenever you add a string attribute, that string value gets parsed as JSON if it starts with "{" or "[". Which means that parentnode.put("request","4,5") would add the string attribute "request":"4,5" to the parent node, the slight variation parentnode.put("request","[4,5]") would generate the attribute "request":[ 4 , 5]. This is fundamentally different, and challenges any assumption the recipient had about types in marshalled data, as the types varies depending on the contents of the strings being marshalled.

Needless to say, I was unhappy. At least by the point that unhappiness was reached, there were now some tests to work out what was going wrong. Which made fixing it possible, the fix being to always single quote strings being added, such as parentnode.put("request","'[4,5]'"). When the put() operation is invoked, the single quotes are stripped, and the unparsed inner values become the attribute. With careful wrapping of the put operations -and ensuring that only one set of double quotes is placed around every string attribute, the tests passed, everything worked, and a new nightly release went out with the issue marked as closed.

Except the next day, the issues came back. Because it wasn't fixed. Not once you took that node, parentnode, and added it under another node: messagenode.put("payload", parentnode).

Doing that appears to trigger a reparse of every string value. Which means all safely planted array strings end up being reparsed, and converted from strings to JSON arrays.

At this point, the unhappiness level changed from "medium" to excessive. With the new tests replicating this behaviour, and no obvious in-source switch to say "be less helpful", the only solution that seemed 100% likely to work with all possible payloads and orderings of payload construction was to rip out the entire library and replace it with one that did not exhibit the same behaviour.

Which is what I spent Thursday and Friday doing: a complete removal of all uses of json-lib and replacement with Jackson. Which, I not only regret in time wasted, but in the way that the json-lib Java-friendly object model is better than the Jackson "look like the DOM" world view, because the DOM, ubiquitous as it is, is pretty painful. Yes, with experience of the XML parsing world, I can certainly use it -but that did not mean I enjoy it. I've had to rip out and place server side and client side code, other things that are visible to other bits of the system that used the same JSON library. I haven't fixed all that code -but add enough downconvert/upconvert that the code compiles and appears to work.

Four hours over two days writing tests to show this problem existed -then
two full days of repair: switching libraries, backtracking on method calls, changing types and seeing won't compile, running tests. Some extra tests this weekend and then only the merge with two days worth of other peoples changes and some reruns and it's ready to commit. That will leave Jenkins to do the final retest and email, then everyone downstream will have to fix their side of the down- and up- converted code (the conversion methods marked as @Deprecated to make them easy to spot), which will take 1+ hour on monday for a couple of people.

When we look back at this, what will we be able to say?

We can now reliably send strings with square brackets around.

This is one of those weeks that I would never use in a motivational talk for anyone interested in taking software engineering.

[Artwork: 3Dom, "Remember the Future", Richmond Road, Montpelier]