Bill Barth - Slashdot User

Comment Re:Hardly A New Problem (Score 1) 112

by Bill Barth on Friday November 23, 2012 @12:22PM (#42074439) Attached to: Supercomputers' Growing Resilience Problems

Suppose that your job is computing along using about 3/4ths of the memory on node 146 (and every other node in your job) when that node's 4th DIMM dies and the whole node hangs or powers off. Where should the data that was in the memory of node 146 come from in order to be migrated to node 187?

There are a couple of usual options: 1) a checkpoint file written to disk earlier, 2) a hot spare node that held the same data but wasn't fully participating in the process.

Option 2 basically means that you throw away half of your compute resources in case there is a failure. Basically no computations being done today for scientific research are valuable enough to warrant this approach. Some version of Option 1 (periodic checkpointing and restarting) is always more cost effective. These systems, in the US at least, are generally between 2 and 10 times over requested by the science community. Taking half away to seamlessly prevent an occasional job death just isn't worth the lost opportunity to more fully utilize the resources.

Option 1 implies taking some time away from your job to do the checkpointing. The vast majority of the time, some sort of OS-level automated checkpointing would be overkill as well. The author of the code knows better when is a good time to checkpoint and when it's a bad idea. I.e. you might consider checkpointing at a phase of the calculation when the data volume required to restore the state is at a minimum even if that means losing some part of a future calculation. Generally the calculation is cheap to redo since checkpoints of large volumes of data are expensive.

In addition, OS-level checkpointing is a hard problem. E.g. if there are messages in flight on the network, do you try to log them and be able to restart them, or do you only checkpoint when the network is quiet? If the network is never quiet on all the nodes in the job, do you throw in needless synchronization that could ruin the parallel efficiency of the job in order to find a place to do your automated checkpoint? If you decide to log instead, where do you write the log data in order to avoid catastrophic failure of each node, and what's the cost of doing it?

If these were just a bunch of VMs running a LAMP stack, this wouldn't be a hard problem. That's basically solved already. Migrating tasks for HPC jobs is truly a hard problem with tradeoffs to be considered.

Comment Re:So (Score 1) 308

by Bill Barth on Friday October 12, 2012 @10:20AM (#41630871) Attached to: Linux Foundation Offers Solution for UEFI Secure Boot

Which is great unless you have 5000 nodes that you need to PXE.

Comment Re:I find this depressing (Score 1) 275

by Bill Barth on Monday August 06, 2012 @01:12AM (#40891619) Attached to: Neutrino-Powered Financial Trading In Our Future?

It hasn't really done much for HPC. It has made a dent in low-latency Ethernet, but that doesn't get much use in high-end HPC (that's IB and the Cray and BlueGene networks, FWIW).

Comment Re:Similar work exists (Score 1) 25

by Bill Barth on Monday July 02, 2012 @06:44PM (#40522227) Attached to: UK Universities Launch Cloud Supercomputer For Hire

It's nothing like TG. TG systems basically gave all their cycles away for free through the work of the Resource Allocation Committee--a peer-review body that met quarterly to review proposals and give out allocations of time. This work continues through the XD program under the auspices of XSEDE.

Comment Re:There is not even a way to remove it! (Score 5, Informative) 346

by Bill Barth on Tuesday June 26, 2012 @08:56AM (#40451221) Attached to: Facebook Says Your Email Is @Facebook

You can't get rid of the address, but you can make it so that no one sees it. You can also display to whomever you like whatever address you like. The settings updates you have to make are pretty straightforward.

Comment Re:Impressive engineering feat (Score 1) 118

by Bill Barth on Monday June 25, 2012 @10:39PM (#40448047) Attached to: Gamera II Team Smashes Previous Best Human-Powered Helicopter Flight Time

The contest doesn't require control, and given the power requirements, it's unlikely that even the best cyclists will every fly one of these around the countryside.

Comment Re:battery life (Score 3, Interesting) 97

by Bill Barth on Sunday June 10, 2012 @02:13PM (#40276185) Attached to: Linaro Tweaks Speed Up Android, By Up To 100 Percent

That's not guaranteed at all. The power consumption of a CPU is a function of a huge variety of things. It's possible that while the run time is shorter, the power draw is higher--possibly more than proportionally higher.

Comment Re:Clarification, as I live here and study there. (Score 2) 386

by Bill Barth on Sunday June 10, 2012 @10:55AM (#40274603) Attached to: RMS Robbed of Passport and Other Belongings In Argentina

Nobody checks your ID when you go to class in the US either, though there's much less of a culture of people just showing up listening in. It would often, but not always, be easier to detect a stranger in a class here, though there are plenty of 500-person freshman biology lectures, too. Typical classes have ~30 people in them.

Comment Re:Go has some good ideas (Score 1) 186

by Bill Barth on Thursday March 29, 2012 @05:17PM (#39515551) Attached to: Go Version 1 Released

You know that Emacs has been able to let you use the tab key to insert the correct number of spaces for the current line for probably 30 years, right? I suspect that Vim can do it, too.

Comment Re:56 gigabit InfiniBand (Score 1) 55

by Bill Barth on Friday September 23, 2011 @05:22PM (#37496446) Attached to: 10-Petaflops Supercomputer Being Built For Open Science Community

The 36-port part is the ASIC. The switch boxes have a lot more ports.

Comment Re:Why did IBM do this, and what next for NCSA? (Score 1) 76

by Bill Barth on Tuesday August 09, 2011 @10:42AM (#37032652) Attached to: NCSA and IBM Part Ways Over Blue Waters

Not that it changes your argument, but you should know that NCSA has a brand new Altix.

Comment Re:Typical (Score 1) 76

by Bill Barth on Tuesday August 09, 2011 @10:28AM (#37032530) Attached to: NCSA and IBM Part Ways Over Blue Waters

It appears to be the latter. The spec is available here. NCSA negotiated a system with IBM, proposed it to NSF under the above linked RFP, went through a peer-reviewed awards process, negotiated an award with NSF, and started working on the delivery and other aspects with IBM and NCSA's other partners. Something went wrong in the last several months, and IBM's pull out was the result. I doubt that there is any more money to be found, and all parties knew what was asked of them in order for the project to be successful.

Comment Re:"High level" programming environment? Sigh. (Score 1) 76

by Bill Barth on Tuesday May 24, 2011 @09:56PM (#36235222) Attached to: Cray Unveils Its First GPU Supercomputer

Have you tried it off-node?

Comment Re:Rack density (Score 1) 213

by Bill Barth on Wednesday April 13, 2011 @10:09AM (#35807364) Attached to: A Closer Look At Immersion Cooling For the Data Center

Have you tried fitting a water block in a blade lately? How about 2000 of them? :)

Comment Re:Rack density (Score 1) 213

by Bill Barth on Wednesday April 13, 2011 @08:23AM (#35806206) Attached to: A Closer Look At Immersion Cooling For the Data Center

Given that you can lay two of these racks back to back and then run them end to end, and that you can remove most of the regular AC equipment from your room, the amount of stuff you can get in your datacenter is the same.

Slashdot Top Deals