Hacker News new | ask | show | jobs
by old-gregg 4387 days ago
I've been SL customer and have using that prior to joining Rackspace and building OnMetal. The point isn't about "lets get rid of hypervisor", the point was to:

* Lets provision as quickly as VMs (as opposed to 1hr+ on SL)

* Lets engineer hardware to deliver maximum uptime via no-moving-parts design (as opposed to vanilla SuperMicro on SL)

* Lets design hardware to deliver maximum unit of work per dollar (like DB transactions/second per dollar, or requests/second/dollar) as opposed to average value elsewhere.

* Deriving from the above, lets just give RAM away nearly for free, and put dual 10Gig network in place, because modern apps should be mostly RAM-based.

* Lets adopt standard OpenStack provisioning API, with myriad of pre-existing tools, community and ecosystem (like auto-scaling, orchestration, etc) as opposed to proprietary API.

The end result is a completely different infrastructure, something akin to what OpenCompute pioneers use internally. This is how running at scale should be like.

As always, I encourage skeptics to spend more than 10 seconds on a product page, because most of the time there're humans behind it, and - in this case for sure - they are way too ambitious to be spending their lives simply cloning old designs.

Enjoy OnMetal, dear jberekw, it's built for critical thinkers (and skeptics! :-) like you and it's awesome - it's going to rock your world.

2 comments

The pricing (on the compute flavor especially) is appealing -- can certainly see uses for it.

Looking forward to OnMetal being available in other regions, especially ORD.

I'm one of the skeptics -- but that's because I've had very few problems with Rackspace cloud instance performance in general; rather, it's the control plane reliability which has caused 99% of the pain.

As an aside, I find the Cloud Servers SLA disappointing, as it defines control plane availability as 1 – (Total API Errors)/(Total Valid API Requests) . Can't launch instances on Super Bowl Sunday? Tough. The failed POSTs to /servers will never trigger a violation, so long as GET /servers is working:

http://www.rackspace.com/information/legal/cloud/sla

Most SLAs remedies are weak ("we break the SLA you get a free t-shirt"), but this one is broken to the point of glossing over fundamental outages.

Agreed -- SLAs are a pet peeve of mine. I wrote a post [0] about this a few years ago, it still feels fairly current. Basically, I don't think SLAs are a useful tool for managing service provider relationships; it's too difficult to capture nuance in a legal document. Transparency and track records seem like a more realistic approach. We've been seeing some movement here, but not much. Most service providers offer very little hard data about historical performance, and don't say enough about their internal operations to let you make a realistic assessment of disaster probabilities.

[0] http://amistrongeryet.blogspot.com/2011/04/slas-considered-h...

We're strongly considering moving our infrastructure to IAD from ORD - we have 60+ VMs in ORD, and their datacenter is at capacity, to the point that only existing customers can create ORD VMs (and there was that instance a few months back where they couldn't provision new SSD block storage).

I know from their sales that ORD isn't on the OnMetal roadmap for at least 12 months (if at all - I believe "it doesn't appear on our 2014 or 2015 roadmap").

> * Lets provision as quickly as VMs (as opposed to 1hr+ on SL)

How many Rackspace customers need a full machine in under an hour?

> * Lets engineer hardware to deliver maximum uptime via no-moving-parts design (as opposed to vanilla SuperMicro on SL)

So just replace spinning disk with SSDs. Unless you're replacing the last moving parts (CPU, PSU, Chassis Fans) with something solid state.

> * Lets design hardware to deliver maximum unit of work per dollar (like DB transactions/second per dollar, or requests/second/dollar) as opposed to average value elsewhere.

This is fair for RAM intensive workloads. Everyone else is already giving you instance SSD access.

> * Deriving from the above, lets just give RAM away nearly for free, and put dual 10Gig network in place, because modern apps should be mostly RAM-based.

Again, perfect for RAM intensive workloads.

> * Lets adopt standard OpenStack provisioning API, with myriad of pre-existing tools, community and ecosystem (like auto-scaling, orchestration, etc) as opposed to proprietary API.

Also another fair point. OpenStack (and its open platform design) is all Rackspace has to compete against AWS and Google.

I don't want to say OnMetal is their "Hail Mary", but Rackspace is exploring their options in the marketplace with regards to an acquisition: http://www.bloomberg.com/news/2014-05-15/rackspace-hires-mor...

To answer your questions:

1. Anybody with $10K+ monthly hosting spend would love to get a "full server" in under an hour. Actually sub-second provisioning would be nice too.

2. You are right, and yes, we've moved power and cooling away from the servers to externally serviceable redundant arrangement. They're truly no-moving-parts. And we've put something way better than SSDs into them.