Saturday, April 9, 2011

Rebuilding Radio NZ - Part 2: The Birth of ELF

In this second part I'll be talking about the birth of ELF.

Warning

Even though its a drag, I'm repeating this disclaimer.

There is no such thing as instant pudding. You cannot copy what someone else does and get the same result.

This post is about a specific site with its own special functional requirements and traffic loads. radionz.co.nz is a public broadcaster's website that includes news, audio and content related to on-air programmes. Traffic loads are very peaky (and high).

This series of posts should NOT taken as advice for or against any particular system. It deals with our specific pain-points and how we are solving them.

You should do your own research and assessment before choosing any CMS or development framework. A good starting point is Strategic Content Management on A List Apart.

Some Management Theory

The manager is responsible for the system in which his staff work. By system, I mean all aspects of the job that contribute to whatever you are producing. The system includes workspaces, office layout, tools, technology, processes and procedures to name a few components.

It is the manager's responsibility to improve the system. In doing so he must understand the difference between problems which are part of the system (built in), and those that are outliers (from outside).

For example, for knowledge workers their computer is part of the system. No one can be productive if their computer keeps failing, is underpowered or does not have the software they need to do their job.

A one-off power cut that stops people working for a day is probably an outlier that needs special attention. (Or may need no attention at all).

The system itself (and everything in it) needs to be designed and maintained. There is nothing worse than a free-running system where components essentially design themselves, are become sub-optimised, failing to work together as a whole. It is very common for processes to become run-down over time and no longer be fit-for-purpose.

The aim is to have stable, predictable processes where you can be sure that content moving through the system meets quality expectations when it is finally published. Efficiency, and replicability are just two aspects of the equation.

The tools that are used to produce and manage web content play a critical role in the system, and one of my roles is to make sure the tools do not get in the way of creating our content.

It is from that base that we considered the suitability of our current CMS tools.

Cracks in the walls

The Radio NZ website was built from scratch - when we started we had no existing processes to support publishing large amounts of web content, and no web infrastructure. We designed new publishing processes and chose our tools (Matrix and a number of custom scripts) based on those processes (I'll be documenting these in later posts).

These processes have been improved iteratively over time. Some of these changes were facilitated by new features in Matrix, others from internal rearrangement. As well as process improvement, we continued to add new content and functionality to the site.

But from late 2009 we found it increasingly difficult to innovate. The modular approach to building sites in Matrix - the very paradigm that got us off the ground so fast - was slowing us down.

Matrix makes The Hard Things simple. Start a new site, set up a home page, a 404 page; all done is 5 minutes. Change content on an About Us page; 1 minute. Setup a form for people to submit queries; done in 10. Display the same content in three different places, auto-generate menu structures; more complex, but still relatively fast to implement.

But for us, some Simple Things were getting harder to do. We were having to create increasingly complex configurations to optimise the display of our content, and create new ways of viewing it. (Examples of this in subsequent posts).

This was largely because our content was stored as individual assets, rather than as modelled data. Each asset knows nothing about any other asset. For example, an audio asset does not know what programme it was broadcast on. A programme asset (a programme's home page) does not know who the hosts of the programme are. And so on.

Some of the asset structures required to support certain features require huge amounts of work to implement.

On top of this was system performance. We are a media site with fast-changing content and high performance demands, and I think the only Matrix customer using the system in this way.

Many pages (like our old home page) were built from many pieces, putting a high load in the system when they had to be rebuilt and cached. With frequent publishing we had to expire the cache as often as 10 times an hour.
In order to deal with our high traffic load it was suggested that a custom caching regime be considered. This would allow us to publish updates 5-10 times an hour, and for the content to be recached more efficiently.

We had already made changes to the operation of the cache (see this old post), and they'd been running for several years, so I had a very good understanding of how this part of the system worked. It was unlikely that these new changes would be of use to other Matrix users and would not become a part of the core product; if implemented, they would be our responsibility to maintain.

The cost of working-around these two problems (asset modeling and caching) - problems that may not exist with other systems - was deemed too high. Sadly, matrix was no longer a good fit for our content or our traffic profile. It was time to consider alternatives.

The decision to change was entirely pragmatic and based on changing business requirements. It was a difficult decision to make, especially after a long history with one product.

ELF is born

Looking at our content, and the sort of features we wanted, it was pretty obvious that a lot of custom code would have to be written.

Very few of our pages are the standard 'edit, upload a photo, update the title' type of content. With this in mind I thought it better to have complete control over all the software, rather than bolt 95% of what we wanted onto an existing product.

Rails looked like a good platform to model and deliver content like ours, and had an excellent local (Wellington) community. There are many development houses and government agencies working with Rails.

So Ruby on Rails it was.

An additional factor was the use of the framework on our company intranet. We had developed a number of powerful modules that could be leveraged for the public website. (In practice, I think we saved about 6 weeks time by recycling existing code).

The name ELF was chosen after a brain-storming session. ELF stands for Eight Legged Freak (i.e. a spider). It was chosen because a spider lives on the web, and because an Elf has 'magical powers' that benefit its users.

In my next post I'll talk about planning the migration of content and the first section we built and made public: Recipes.

Rebuilding Radio NZ - Part 1

This is the first in a series of posts explaining how (and why) we are rebuilding www.radionz.co.nz. I'll be examining the technology behind it and looking at some of the difficult choices we've made along the way.

Warning

This post is about a specific site with its own special functional requirements and traffic loads. radionz.co.nz is a public broadcaster's website that includes news, audio and content related to on-air programmes. Traffic loads are very peaky (and high).

This series of posts should NOT taken as advice for or against any particular system. It deals with our specific pain-points and how we are solving them.

You should do your own research and assessment before choosing any CMS or development framework. A good starting point is Strategic Content Management on A List Apart.

Beginnings

Since October 2005 the site has been running on MySource Matrix (now Squiz Matrix). We started the build of the site around Easter that year, meaning we've used the system for six years - not a bad life for any piece of software. We know it very well.

Matrix was chosen after an exhaustive process where we evaluated dozens of web CMSs and called for a Request for Proposal (we got over 40 responses). The project took a year to complete as we had no existing infrastructure or business processes to support the content we wanted to publish.

We ran the site in-house for 3 months to bed in new publishing processes and iron out any bugs.

The primary reasons we chose Matrix was depth and breadth of functionality and the ability to build sites without needing a programming resource.

The previous version of the site was based on a custom built CMS (PHP), and the requirement to have access to a programmer was a constraint we wanted to avoid for the next version of the site. The only custom work was a Matrix asset to support audio content.

With Matrix we were able to quickly build almost anything we could conceive and have it live quite quickly. We also had the ability to try stuff out, modify it, and then release, all without a code editor in sight.

In five years we grew the site from having a rolling one-week back-catalogue of audio content for some programmes, to having over three years of back-content for most programmes. Traffic increased 10-fold.

In late 2009 we started to experience some pain, and by early 2010 we made the decision to move on from Matrix. This was not a decision made quickly or lightly, and it was based on deeply pragmatic reasons that I'll explain in future posts.

The first major reason was increasing difficulty in managing our content - at the time we had about 5,500 individual programme pages (today it's about 7,000). Moving around the site between pages, and the time required to update content was limiting the amount of content we could publish and causing frustration for editors.

The second was performance. We have a lot of content that is updated frequently - a fast moving news story could easily be updated 10 times in a morning - and the system was not able to cope with our requirement to refresh the caches for these pages for every update (more detail on this in part 3), at least not with the hardware we had at our disposal.

You could say that we outgrew the system. Not because there was necessarily anything wrong with it per se, but because our operation had grown in a direction where the system was no longer a good fit. Pragmatic, as I said.

Today (April 2011) our site is running partly in Matrix and partly in the new system (called ELF). This has presented some challenges both in integration (keeping the experience seamless for visitors), migration of content, and in the training of content editors. I'll cover all this in a later post.

In the next post I'll talk about the birth of ELF.

Saturday, April 10, 2010

Showing hash differences in Ruby tests

Today I was writing some tests for a Rails plugin I'm working on. The test had to compare a hash output by a method with an expected hash.

This is reasonably standard stuff - when assertion_equal failed it printed out the expected and provided data. The problem in my case was that the hashes have about 50 elements, and I only want the difference to be shown in the test output.

I knocked up this function to solve the problem, and included it in my test class:

def assert_hashes_equal(expected, actual, message=nil)
full_message = build_message(message, "Hashes were not equal, diff was:\n.\n", expected.diff(actual))
assert_block(full_message) { expected == actual }
end

So now I get this:

Hashes were not equal, diff was:
<{:plays_in_category=>"2", :dynamic_end=>"MFT", :percent_back=>"100w"}>.


Note that Hash.diff is part of Rails.

Here is what the output looks like in practice:

1) Failure:
test: SelectorImporterTest Concert Selector Parse Song 0 XML should Hashes should be equal. (SelectorImporterTest)
[/test/unit/selector_importer_test.rb:500:in `assert_hashes_equal'
/test/unit/selector_importer_test.rb:368:in `__bind_1271217771_635209'
vendor/gems/thoughtbot-shoulda-2.10.1/lib/shoulda/context.rb:253:in `call'
vendor/gems/thoughtbot-shoulda-2.10.1/lib/shoulda/context.rb:253:in `test: SelectorImporterTest Concert Selector Parse Song 0 XML should Hashes should be equal. ']:
Hashes were not equal, diff was:
<{:mood=>"3"}>.

Saturday, March 20, 2010

The New Accessibility Wave

There is a new accessibility imperative.

You may have heard of it: Mobile.

The mobile user experience is fundamentally different from browsing the web on a desktop grade computer and a standard web browser.

Touch devices do not have a keyboard. You cannot control + click. You cannot right click. You cannot click and drag items. (Usually). Non-touch platforms have other challenges. Some do not support certain plugins.

This is going to require a reworking of the way we build websites.

Design and functionality will have to be layered onto content and structure based on the consuming platform.

The techniques to do this already exist. One is progressive enhancement, but at the moment this is more about adding advanced features for desktop-grade browsers.

It needs to go further than this. Websites have to work for everyone without discrimination. If they don't work, the user's experience of your site will be poor, or non-existent.

The reason I say 'new' accessibility wave is that we've already had to make websites for technology that interacts with pages differently from desktop web-browsers: screen readers.

Many ignored this old wave completely, deeming that market as irrelevant, insignificant or uneconomic.

I have bad news for you. That wave is being joined by a new one - tens of millions of people using non-traditional browsing technology, and growing every day.

If you've not thought about it, it is time to think about inclusive websites again. Websites will have to change to support not only screen readers, but mobile, TV and who knows what else.

Open technologies are going to be a big part of this - HTML5, CSS and Javascript.

If you didn't catch the first wave, you have a lot of catching up to do.

Friday, October 9, 2009

How to Remember CSS Shorcut Order

I have always had trouble remembering the correct order for CSS shortcuts like this:

margin: 3px 4px 2px 3px;

The order is Top, Right, Bottom, Left.

I have just found two ways to remember this.

The first is using the word trouble:

T R o u B L e

The consonants give the order.

The second (pointed out by a colleague this morning) is using a clockface starting at the top and going clockwise.

12 is at the Top
3 is on the Right side
6 is at the Bottom
9 is on the Left.