Wednesday, September 26, 2007

svn.perl.org password/account management

For a long time we had a confusing thing where to get a svn.perl.org subversion (and groups.pm.org webdav) account you had to go to auth.perl.org and set it up there.



We've improved it a bit now so resetting your password (or just setting up the account initially) you just go to svn.perl.org. Whee!



Let me know (ask@perl.org) if you have any trouble with the new site.



Tuesday, September 18, 2007

Week of failures!

Geez! This week it seems like everything is falling apart over here. Servers that have been running for hundreds of days are acting up. Configuration files last changed 3 years ago needs tweaking. Programs not changed for 4 years have bugs popping up.



We had an outage of cpan.org mail (mail bounced) early this morning and later in the afternoon many of the perl.org sites were unavailable for a few hours. One of our database servers got "stuck" and it had a cascading effect on some of the web services.



Apologies! IM me at xmpp:ask@plys.net if you see anything else amiss and we'll get it fixed in a hurry.



Monday, September 10, 2007

CPAN Search adds Gravatar icons

At the request of Michael G Schwern gravatar icons have been added to author pages on CPAN search.



Any author can get a gravatar icon on their page by registering at gravatar.com using their CPAN email alias.



Wednesday, July 18, 2007

svn.perl.org outage

The svn.perl.org server went kaboom. The kernel killed itself with lots of Out Of Memory errors.



I can't figure out the magic incantation to get into the Lights Out Management on that box (it's an old IBM server) and it appears we set it up to use that rather than the Cyclades console and power management.



Anyway - the datacenter staff has been sent to our cabinet to get it booted, so hopefully it won't be too long before it's back. Worst case it'll be up tomorrow morning (PST) when Robert or I go down there and find out what's going on (sorry, you will have to do without the svn server for 7 hours).



Thursday, July 5, 2007

Thanks Leon!

We just wanted to thank Leon for his assistance in providing a copy of BackPan we could use to validate our local copy against.  Be sure to check out his recipies!



Sunday, June 10, 2007

~All is well

Almost everything is back up. Apologies for the super-extended outage. I suppose it's bound to happen every few years, but we'll get to work on making sure this particular failure won't happen again.



Roberts been nudging the MySQL databases and no data should be lost. (A few databases did disappear, but we should be able to find a backup or otherwise restore it from exported data for all of them). Cross your fingers for us!



We're still working on making sure all the RT data is in good shape before putting it online, but it should be back one of the next few days.



Let us know at ask@ and robert@ if you see anything strange in the next few days. (Or when RT is back then at webmaster@ ...)



Thank you for your patience!



Friday, June 8, 2007

RAID went boom part 2

The system is still recovering from the failure. Details are long and boring, but computers suck. :-)



Ironically we've just been talking (again) about getting the services running on the failed server made redundant, but haven't actually done it yet. It's our biggest SPOF by far. Grrh.



Right now the ETA is "in the morning" (PST). Robert is sleeping and will check on it when he gets up. I'll keep an eye on it for a little bit longer (and then go to bed because tomorrow it's my birthday and hopefully I'll have to eat a big cake or something!)



Something Go Boom

Something's broken on our main fileserver.  It's related to our disk array.  We're working on fixing it.  More news as it happens.  Lots of stuff may be broken until we get it back up.



Update 7:10pm:  We've had a disk failure, and the raid volume got unhappy.  We don't think we've lost any important data, but are verifying.  Things are going slow because it takes a while to check half a terabyte of data.\



Update 11:40pm: This is taking a while to resync and fsck.  (Plus we accidentally aborted it halfway through.)   We're going to go to bed, we'll try and have services back in the morning.



Thursday, May 24, 2007

May is for Maintenance

Friday, May 25th at around 10:30am PDT we're going to begin performing maintenance on some of our servers. For a few hours you may notice some oddness or unavailability. We'll post here with any updates and when we're done.



Update:

Maintenance is over, it went off without a hitch. RAM upgraded, bad disks replaced.



Monday, May 21, 2007

subversion sidegrade

You may have noticed we had about two hours of subversion downtime tonight on svn.perl.org. You shouldn't (we hope!) notice anything different, but we've converted our repositories from the BerkeleyDB backing store to FSFS. (Design Doc). BDB has served us well with no major issues for years, but it was time to change.



Wednesday, April 4, 2007

RHEL5

We've started upgrading some of our boxes to Red Hat Enterprise Linux 5 which includes virtualization. We haven't entirely figured out how we are going to use it, but it'll likely make it easier for us to host more things, which is good...



Thursday, March 22, 2007

Bad Spam Day, But Nice Weather

<me> Today is a bad spam day.

<someone> nice weather. all the spammers are coming out.

<someone else> no comment

<someone> crap, are you spamming today?



There's a lot of email spam slipping through our filters over the past few days, and that stinks! (It's getting through other people's filters too - so we're not the only ones having trouble.) We're working on tightening things up - but we have day jobs too - so please bear with us and don't report us to spamcop.



Monday, March 19, 2007

New CPAN Search mirror(s)

One of the boxes serving CPAN Search to the US and the world has been about to fall over for a while (one of the disks hangs the system for a minute once in a while with scsi errors...).   We are working on setting up new a pair of load balanced (for redundancy) mirrors provided by YellowBot.



Ideally we'll get it setup so the DNS servers will check and notice when one of the mirrors are out too... (patches to pgeodns are welcome - we even have a wiki for it and other perl.org infrastructure things, send an email to get invited to it).



We could still use a well-connected box for CPAN Search in Asia -- it probably wouldn't get much traffic, but it'd decrease latency for users there.  Minimal specs required (it can be in a virtual box, vmware or xen): ~2.5GHz CPU, ~2GB ram, not much disk space, ability to install RHEL (we'll provide a license).