Thursday, 26 August 2010

Cool!

Some years ago, we undertook a small experiment with our server room. We had heard that other people were reducing the amount of A/C cooling they used and we wanted to see if it was appropriate for us.

Like a lot of other places, our small server room was kept cool to keep the servers cool; if we were to spend any length of time in there, we would need to put on a jumper or even a fleece to stay warm, as the room was around 10 degrees centigrade. The A/C units were running non stop, and we wanted to see if we could reduce the electric we used.

Essentially, we made a load of measurements to get a base line. These included the core temperature of the UPS, some measurements of the servers and various places within the server room. We were fortunate that our engineering manager had a device that we could borrow for this as he was conducting a number of tests to help the company work towards ISO14001.

He also had a device that allowed us to measure the amount of power drawn by various devices – we seemed to get a couple of slightly odd readings, but when we discounted those, the average values appeared to match what would be expected. We therefore assumed that the errors we had were down to incorrect use.

Having got our base line values, we then started to increase the ambient temperature of the room, and examine what affect this had. Each time, we would leave the changed settings for a couple of weeks to see what would happen; in each case, there was no sign of distress on the servers, so we were able to increase the temperature again.

After some time, we found that the “sweet spot” was between 20 and 24 degrees centigrade. Above 24 degrees, we would see the fans in the servers starting to work much harder and draw more power. Below 20 degrees, the A/C was still running almost all the time. However, in that range, we found that we had the A/C unit running at its least power draw whilst the servers ran at a comfortable level.

We found that in the racks, we had a few “hot spots”; places where the temperature was quite a bit higher than the ambient temperature of the room. We were told that this is normal and generally considered a good thing; these create a thermal current that allows the cooling to happen naturally. The interesting thing was that although the ambient temperature increased by 12 degrees, the hot spots only increased by 3-4 degrees.

Part of the work meant that we had to make sure that the racks were properly positioned in the room to allow for adequate air flow, and the direction of air from the A/C also had to be optimised to prevent “air curtains” forming at various places. We also had to make sure that things such as blanking plates were used to ensure a properly controlled air current within the racks.

Although this all sounds very grand, the room is quite small and most of the work was done in between our normal activities. We were able to make use of some additional advice from the A/C supplier, but that was relatively minor. The total amount of work required was actually quite small, but the results have been very good. We have seen a reduction in power consumption of just under 50% for the server room as a whole – which translates into significant cost savings.

I’ve added a link to a resource that I would recommend to anyone wanting to do work on their server room facilities. It is primarily aimed at North America, but there are some bits that are specifically for the European market. It will take some time to go through all of it, but I consider that it would be time well spent.

http://www.schneider-electric.com/sites/corporate/en/products-services/training/energy-university/energy-university.page?tsk=77518T&pc=26947T&keycode=77518T&promocode=26947T&promo_key=26947T

The really good thing - we now have a server room that we can work in, in reasonable comfort all the time!

Tuesday, 24 August 2010

I'm back!

It's been a while since I posted anything; 6 months in fact. It's not a case of having nothing to write about, far from it. I've just been very busy, plus I've been a bit more active in other areas.

One thing that I thought would be appropriate to point out is a Microsoft resource at: http://www.microsoft.com/uk/business/peopleready/technology/ioassessment/osyci/survey.mspx
This allows you to take a "survey" that can give you an indication of the status of your IT provision. I first came across this a while back and I found it very useful as part of the planning process. In order for you to reach a particular destination, it helps to know where you are starting from, so you can use the right directions.

Essentially, Microsoft suggest that IT departments can be classified into one of 4 levels based upon standard practice. Five years ago, we would have definitely been classed as being at the lowest level, "reactive". The IT provision was based around fixing problems after they occurred and very little thought went into planning or preparation.

We've slowly moved through the various stages, going from "standardised" to "rationalised", and are now pretty much at the top level, "strategic". There are still a few areas that we could improve upon, but that will always be the case. However, the IT is now a solid platform that people can use. We don't get the network failures, system crashes, or data losses that used to occur. Resources are there and available 24 x 365 for people to use, and generally they can access them using whatever device is appropriate.

Now although this all sounds great, there is unfortunately a fly in the ointment. The biggest problem is still the unit that is positioned between the chair and keyboard! It has been identified that we need to get people better trained, but somehow that never seems to get translated into action. Once of the worst instances was of a person that had been with the company for some 8 years. Unable to logon, the person phoned the helpdesk to ask what her user name was! (She normally didn't have to type that in, as it just appeared in the login box.)

I would encourage everyone to take a look at the Microsoft Core Infrastructure Optimisation resource. I think that you'll find it of significant value and help.

Friday, 12 February 2010

BCS - Computer Forensics

For some time, I've been working towards a post graduate degree through the Open University. It's hard work, particular after a long day when all you want to do is switch off and relax. However, I find the courses fascinating and of help to me in my daily work, so I keep on working on it.

The last course I did was particularly interesting - Computer Forensics and Digital Examinations. This is a very technical issue, but it also requires an understanding of legal procedures. It isn't enough to say "I found so and so", you have to demonstrate that the evidence is relevant, accurate, consistent and to present it in a way that non-technical people can understand it. I found it all really interesting, if not totally linked to my daily job.

So when the BCS South West indicated that they were holding an evening event and the topic was Computer Forensics, I jumped at the chance to attend. It was at the University of Plymouth, which is a really nice venue, if a little bit of a trek to get to from where I live. The speaker was a visiting professor, John Haggerty from Manchester and the presentation was lively and informative. The actual notes should be available at this link. http://www.bcssouthwest.org.uk/server.asp?page=pastevents

For me, the presentation covered most of the items that I has previously studied and it was really good to refresh my memory. It was also interesting to see that after such a short time since I did my course, there are a number of changes that have occured and the discussion after the talk highlighted some of the issues facing practitioners in that field.

One thing that is of interest - Digital Forensics is a field that is wide open for people to move into. However, there are a lot of people that think just because they have a small amount of experience in running a computer, they think they know what to do to examine it. Professor Haggerty referred to this as the "CSI" affect - people see the TV shows where someone drives an expensive car, goes to a pristine work space and in half an hour recovers all the require information to solve the case (and the impossibly attractive woman is suitably impressed by the display of brain power!).

In reality, Forensics is a long tedious job. Everything has to be documented, step by step and assumptions made have to be justified. There are a number of practioners that have had their reputations destroyed by a simple mistake, and once that happens, they are unlikely to be able to work in the field again.

As the technology moves on, the process of the examination gets harder - I can remember when I bought my first hard drive of 20 MB and I wondered how I would ever fill it up. I now regularly work with physical hard drives of 500 GB and logical partitions of over 1 TB. To properly analyse and document such a drive can take a very long time and new tools are being developed to try to make the analysis easier, but it still requires considerable patience.

But all in all, a great evening - a fascinating topic, well presented. An dfor those IT people that think the BCS is only for academics, I would strongly suggest that you go along to one of the (free) events - I'm sure that you'll change your mind.

Friday, 29 January 2010

Hard driving SQL

We have been working on installing an SAP ERP system for some time now. It went live in the latter part of 2009, and almost immediately we started to get some performance issues. After some discussions, we were advised that we should move a number of components form the SQL server to separate disks.

The server had originally been set-up to the specific instructions of the system integrators, and they had carried out the installation of their software. We had 2 logical drives; the operating system on the C: drive and the rest of the product on drive D:.

Essentially, they now advise that we should put the paging file, tempdb files, and transaction log files all on separate logical drives. This does make sense; with the extra drives, there will be less data being processed at the same time by the same equipment. However, the server we have is an HP Proliant DL380 with space for just 6 drives. As all the slots are full, we can’t physically add any more to the existing device.

However, there is a way around this; HP sell external disk arrays which can be added to an existing server. In our case, we obtained the MSA 20 unit which hold 12 SATA drives and this is connected to an HP 6400 SmartArray controller card. We ordered all of the required equipment back before Christmas, but unfortunately we had a series of problems getting the hardware. The bad weather didn’t help as we are a bit out on a limb, but the various bits were coming from different depots, so weren’t despatched together.

Laste week after all of the equipment had finally turned up, ee set-out to do a test of the process of adding the hardware and this went through fairly well. It toook us about 5 hours as we wanted to double check everything at each stage to make sure it worked; we had not had the chance to do something like this before and wanted to be certain it would work. We made notes of the steps and waited for the Sunday so that we could make a start on adding the new hardware to the main system.

The controller card was very easy to add. Pop open the cover, lift out the holder, insert the card and replace the cover. I also connected the cable to the disk array at the same time as I found that easier than trying to fiddle about in the back of the rack trying to make the connection. The slot that the cable uses on the back of the card is quite small and difficult to reach when the server is back in place.

When we fired up the server, it ran through the normal POST routine, and it quickly identified the new Smart Array device. It took a while for the disks to initialise; about 12-15 minutes for them all. However, we then hit our first snag; when it reached the end of the initialisation, it suddenly crashed and re-booted. Funny thing though, when the server restarted, it went back to the initialisation routine and then completed perfectly.

It was then necessary to set-up the logical drives and this is really easy to do. Within the configuration utility, just select the physical drives, the type of RAID and away you go. We chose to put 3 drives at a time in a RAID 5 configuration. It should give the space we need, the protection that it wanted and we get 4 logical drives (12 HDD divided by 3 = 4). With all 4 done, we could then re-boot the server, and see the new drives in the disk manager – we set it to create a new partition on the logical drives and format appropriately.

All of this took us about an hour, perhaps just a bit over. We then moved the paging file and set it to a slightly larger size than before – a quick reboot and still everything was going well. We copied the tmpdb folder to a new drive and then used a SQL script that we had found for dropping it and then re-attaching to the new location. It took literally only a few seconds to do and we were starting to get really cocky. Then it all went wrong.

We stopped the SQL service to copy the transaction log over – all 38 GB of it! We then started the copy process and it took ages. It seemed to copy about 8GB and then it would pause for ages (almost 20 minutes), before then carrying on. We got a point where it had reached around 12-14 GB and the damn server blue screened (one of the few occasions that we have seen Windows Server 2003 do a BSOD).

It turned out to be a paging fault error – once started we modified the paging file to put it back to the same minimum size that it had been, although we left it on the same max size. I restarted the copy process and we waited.. and waited… and waited…. and waited…..

Evetually after about another two and half hours, the copy process finished. We then ran the SQL commands to change the database to point to the new trans log location and once done, we verified that this was correct. We then ran up the ERP to make sure that it worked and it was good. By this time, it was well after 1:00 pm – we quickly finished everything off and locked up, then headed off to a local watering hole for Sunday lunch on the company.

And just to finish the story off, the technician’s wife works at that hotel. Whilst we were eating, she sent a note through from the back room, demanding to know where her dinner was. So a small piece was cut off of the dinner and put on a small plate to be sent out to her – 5 minutes later a message came back demanding to know where the ketchup was!

Wednesday, 6 January 2010

New year plans

So the holidays are over and we are all back to work – well almost. Unfortunately, the bad weather has caused some disruption, as a number of staff can’t get into work. Although that hasn’t affected IT staff, we are having to a do few things to help others out. Bet we don’t get any help from them when we need it later in the year!

I like to plan out what work we have to do – preferably at least a few months in advance. As such, I have a list of jobs and priorities against them and this gets updated throughout the year. At the moment, there are a large number of items for the next 3 months and quite a few for the second half of the year.

We are planning to go on a couple of specific training courses, there are some hardware and software upgrades, a couple of events that I feel would be appropriate for myself or my staff to attend and there are a number of jobs that need to be done as part of rolling maintenance programmes. We also have several projects under way and the various steps need to be arranged in the correct sequence and fitted in amongst the other work – plus of course we have the occasional problem that needs to be supported.

Unfortunately, there are several jobs that we cannot yet schedule – we are waiting for information from other people. One of our sites is proving to be a bit too small to handle the work load, so the company are looking at alternative locations. However, the senior managers can’t decide which of the newer sites would be most appropriate, so we can’t yet arrange for any work to be done that is required. Of course we know full well that when they do finally decide, they will expect all of the work to be complete within a few days!

In fact that move is going to be a much bigger task than they anticipate – once the decision is made they will then argue over the layout of the place and almost certainly, will change what they want on a daily basis. We will be cabling up the site for a network ourselves – it saves the company quite a bit of money although it does take up a bit of time. I’ve designed a particular method of network architecture that really works for us, and provides a great deal more flexibility and scalability than the way that these installatins normally get done. Most of the people doing cabling appear to be electrical installers, and they think CAT 5e can be treated like standard 2 core and earth and they seem to have a real problem if you ask for work to be done in a particular way.

On top of that, we have get the telephone lines moved, get an ADSL connection and move all the IT equipment ourselves – the last time we had a move, we also ended moving all the desks and cabinets as well. The staff seemed to think that they could just close down the PCs, put on their coats, pick up their handbags and walk to the new site to find the desks all set up, the PC installed and turned on for them! They were rather upset to find that they were expected to do some of the work themselves!

So January is looking to be quite a busy month, what with one thing and another. Happy new year!

Monday, 14 December 2009

Iiiittttsss Chriiiiissssttttmmmaaaaaassss!

Somone mentioned the old Slade hit from the 70's and I haven't been able to get the damn tune out of my head all day! I think that it's going to drive me crazy! (Mamaa, weer allll crazeeee now!)

Many years ago, on 24th December, I would stay right to the end of the day, and last thing would shut down all of the servers. No-one would be back into work until the first week of January, so it seemed pointless to burn all that power for no reason. Plus it gave the equipment a chance to be shutdown properly and restart. This doesn't always hurt as it can clear out any rubbish in memory.

The trouble was that the CEO felt lost without his email - after we gave him VPN access, he wanted to be able to check his email on Boxing Day, just because he could. Then of course, he wanted to be able to check the sales figure - why? There have been no sales and won't be for 2 weeks - but he wants it, so he gets it. And of course, that means all of the ERP systems have to be running. By the time that you work out which systems he might possibly want, it's easier just to leave them all running. (And of course, you know that he is going to phone up to check if the figures have been updated!)

So we don't shut things down anymore - and that means we have to keep an eye on systems to make sure that nothing untoward is happening. As you can imagine, the WAGS take a dim view of this - it only takes a few minutes to logon and make sure that each of the servers is up and running, but the amount of time is not the issue. We have automated alerts to let us know if specific events occur, but it's not quite the same and there is always a possibility that the relevant alert doesn't get through.

So the laptop is going to be hidden away somewhere, and an excuse made to either "take a nap" or "pop down the pub" - then a quick logon just make sure it's still all OK.

Whatever; we are fast approaching the holidays and the end of yet another year (where does the time go?) From my staff and I, the very best wishes to all the readers of this blog and to all the hardworking IT staff wherever you are. Have a great Christmas and try to enjoy whatever time you are allowed to take off. See you all in 2010!

Tuesday, 1 December 2009

Up in the clouds

One of the hot topics in IT at the moment is “cloud” computing. Effectively, outsourcing your hardware to a dedicated data centre. A lot of people try to convince me that this is the way forward, that everything should be put “on the cloud” and that this will save astonishing amounts of money. I’ve seen some of the calculations and I am not sure that they always stand up to scrutiny.

For example, I looked at a Dell PowerEdge unit – the cost to buy outright (£1,200) was a bit higher than the cost to rent in a data centre for a year (£700), but obviously over a longer period such as 4 years, it would work out cheaper. There is an advantage to the cloud offer in that they would replace the equipment (probably with newer equipment) at a set point, but then it doesn’t appear on the asset ledger in the company accounts, which upsets the beancounters.

Of course the purchase price doesn’t include the Operating System, whereas the cloud offer usually does (but not always); and there is the cost of electric to run the item and to provide cooling which have to be factored into the equation. There is also a need to provide anti-virus protection, patch updates, data backups etc. Again, that is not always included in the price of the hosting contract and so might need to be added to their quoted price – something that is always clear.

In addition, there is the cost of managing the unit – and they don’t always provide all of the management services that might be needed. In most cost comparisons, they show a figure for on-site management (and I sometimes feel that these figures are inflated a bit) - but then they don’t include similar values in the cloud offer even though it would be appropriate to do so, making the comparisons meaningless.

Suppose the 4 year basic cost of renting the server in a data centre would be £2,800 – reading the small print of some hosts, adding in the other items could take it to as much as £4,500. My calculations show the internal cost of the device for keeping it on site could be about the same, perhaps just a little more. Certainly the outsourced system might still be cheaper, but not by that much.

Then there is another point – what happens when things go wrong. It doesn’t happen that often, but when it does, the PTB want to know that someone is working on the problem. They like to be able to go into the server room, and for staff to point out flashing lights, explain what is happening – it gives them enormous comfort to see that someone is on the job and that the problem will be resolved evetually. This can’t happen with an outsourced system – even with numerous phones calls, they just don’t get the same level of reassurance, and you cannot put a price on that.

Now I will accept that I have used very generic figures – and to be blunt, most numbers can be manipulated to show pretty much anything that you want. Ultimately, it should be down to each individual case to be decided on it’s own merits. If it makes sense to keep it in house, then do so; if it is cheaper to host outside then that has to be the right decision.

For example, we have our company websites hosted externally – the cost is far cheaper than we could do it for as we don’t pay for a whole server box, and in addition, we don’t have to provide 24 x 7 support which would really rack up the support cost. However, we maintain our own CRM system – we checked it against SalesForce.com and our internal system works out at half the cost over 2 years. We also maintain our own ERP system – we were offered the chance to have it outsourced, and the cost of the management fees per year alone was more than the wages of our entire IT department.

So I suppose my advice would be to look at the numbers very carefully – make sure that you are really comparing like for like. Then think about the importance of the systems to the business and what would happen if the external system failed and how much of an issue it would be. If the risk is acceptable and the figures check out, then by all means outsource it. But I would strongly suggest that for many people, cloud computing is not the great panacea that it is made out to be, and that it would be appropriate to think carefully before rushing headlong into a situation just because it is the latest, greatest thing.