Showing posts with label AVG. Show all posts
Showing posts with label AVG. Show all posts

Monday, 6 June 2011

BES Hell and the Case of the Missing Window

I was called on Friday... BES Express server down... numskull pulled the plug out of the wall... possibly during a SQL write... pages broken... help!

BESX - version 5.0.2 (Bundle 14)

Where do I start?

Logged into SQL Manager and ran "dbcc checkdb" on all the databases.... consistency errors in some of the tables in BESX... ran "dbcc checkdb (<dbname>, repair_allow_data_loss)" ran "dbcc rebuildindex <tablename>" on the tables... 12 consistency errors fixed...  I accept some data loss in this...

BESX Administration Service - "cannot find the requested page" or "Internet Explorer cannot display the webpage." - the same goes for Web Desktop Manager.  The joy of tying up all of your software's capabilities into a single jvm.dll based web UI... you can't be blamed if it doesn't load up....

So what is to blame?  I don't know.. I search high and low... In the end I uninstalled BESX.  To my regret.

One step back... I wanted to repair install BESX.  So I ran the installer - it stopped - error 2803.  I stopped all the services and tried again.  This time to install BESX "over the top".  I got through to the "Use Existing Configuration Database" .. cool... I can simply install this over the top...

Too soon! "The BlackBerry Configuration Database that you specified is associated with a previous installation of the BlackBerry Enterprise Server or BlackBerry Professional Software. You must create a new BlackBerry Configuration Database for BlackBerry Enterprise Server Express or specify a database that is associated with a previous installation of BlackBerry Enterprise Server Express."

I got fed up with it here - it IS a Config DB for BESX not BES - what are RIM talking about with this error message?... time was running out and I it had to be ready for Monday morning... so I uninstalled BESX.  I regret.

Oh joy... I bet this has something to do with AVG... (lol... I did check but I think AVG was acting normally... it had 58,124 "Virus Found FakeAlert" messages that took forever to not delete.. so when I came back much later a) the program had crashed and b) they had not been deleted ... i so love AVG - not)... 

I found somewhere online mention that I should check that the user BESADMIN has "Allow Log on Locally" as a right before I install BESX (on a side note - I checked through the Release Notes for BESX 5.0.3 (Bundle 12) and decided to download that - in case they'd fixed the 'configuration database is associated with BES and not BESX' error message (bug)... so .. Local Security Policy... er... what's going on here?
MMC could not create the snap-in
MMC could not create the snap-in The snap-in might not have been installed correctly.
Name: Group Policy object Editor
CLSID:{8FC0B734-A0E1-11D1-A7D3-0000F87571E3}
Darn it... something messed up with secpol.msc... Apparently we need gpedit.dll and gptext.dll... they are in System32 folder... and they are registered.. ("regsvr32 gptext.dll" worked but gpedit.dll threw a wobbly...) ... But that didn't solve it... I also need FrameDyn.dll from the wbem folder... it's there... and the program has to find that on the environment variable, PATH

Right... System Properties -> Advanced -> Environment Variables -> PATH= ... er... "...C:\System32;C:\System32\wbem;..."  WTF?!  Where has the "\Window" bit gone?  It goes to show that Windows Server can survive, just about, with this missing ... if you search around online for "c:\system32\wbem" it should not exist ... but it does exist - in one place someone uses it in a solution for the FrameDyn.dll ..  we are pointed to this kb http://support.microsoft.com/?kbid=826282  - be careful with implementing the solution... check that you don't just have a missing 'Window' next to the wbem path.

I wonder now... in retrospect, whether this missing Window was actually the cause of BESX not being able to load the Admin Service or Web Manager.   It can't check the security principals etc... It can't find WMI... 

BESX installed... we couldn't restore the old database over the top.. or copy data from the old db to the new... so we have had to wipe all the users' Blackberrys and reload them... thanks RIM...

The night was so long I can't remember if that was all there was... Space was an issue... an old server, so of course the Windows\installer folder was taking up several gigabytes which can't be removed safely... and the size of the system drive was only 12GB... from the days of small disks...  time for a complete re-install perhaps... 

I doubt anyone who reads this will have exactly the same problem .. this is a Blackberry Enterprise Server Express 5.02 SP1 issue I think.. plus the misfortune of a power failure (or a corrupted BESMgmt SQL DB from any other cause)... together with a malformed PATH variable... you have to be pretty unfortunate to have exactly the same symptoms... :)

I did wonder if it was a recent Windows Update that wrote the PATH variable incorrectly...  it appears to be somewhat common... I don't think people write "C:\System32\wbem" .. but someone has written: "%SYSTEMDRIVE%\system32\wbem" instead of "%SYSTEMROOT%\system32\wbem"

Let me know if that's the work of a virus ... that's the one thing I didn't bother to look into...

Tuesday, 27 April 2010

Blacklisted, Router Spam, Rootkits, Worms

A client has been infected by something nasty... email is not getting through to clients due to blacklisting...

1)  www.mxtoolbox.com  - Lookup the mail server MX record, check for blacklisting.  Find out the date that it happened and any reasons you can find.
2)  Find/fix the infected PC...  As I work remote, I set up a batch file that users can run... 
tasklist /svc > s:\%computername%-tasks.txt
netstat -nabvo > s:\%computername%-netstat.txt
;pause    [ remove commenting semi-colon if you want to]
This assumes there is an S: drive - replace with a network drive letter that you have access to - and save the file as something like 's:\commands.bat'.  Ask someone to run it on all computers.

Check each -netstat.txt file for numerous connections being made where there is a :25 after the Foreign Address (search the text file for ':25').

I've found one PC making such connections...
TCP    10.0.0.12:27163        208.123.68.19:25       SYN_SENT        700
  C:\WINDOWS\System32\mswsock.dll
  C:\WINDOWS\system32\WS2_32.dll
  -- unknown component(s) --
  C:\WINDOWS\system32\kernel32.dll
  [services.exe]
This suggests a dodgy mswsock.dll or software running that is spamming out.  I have checked the other PCs and none of them are infected.  So now to Remote Connect (RDP) into the PC and have a look.

No processes running ... rootkit? .. Ran MalwareBytes' Anti-Malware and found Adware.Starware ... Security Center warning if no Firewall or Antivirus was disabled... AVG looks like it got a random letter named dll file...  and I found Run settings ('rundll32 /dll, startup' in both Local Machine and Local User / Software/Microsoft/Windows/CurrentVersion/Run ... AVG reports the Hiloti... does that send out on smtp?  there's also another file reported as Trojan Horse Generic 17.BEDK.   Adobe monxga32.exe... not Adobe but in a folder called 'Startup' in the Adobe menu group in the Start Menu.

ThePhone.coop has still not responded with a Smart Host for relaying emails... which means their clients will not get their emails... Surely the service assistant didn't need to ask an engineer to call me back just to tell me the name of their Smart Host?

Anti-Malware found 109 issues.. but nothing substantial...  now running AVG's Rootkit Search...  That has not fixed it...  Computer rebooted and still infected.

I have connected to the router and added a new rule to block all SMTP Port 25 traffic from all LAN addresses except the server to any address on the internet.  That will stop the blacklisting organisations from blocking us... and will stop anyone else getting dodgy emails from us.

Now I can remote desktop back into the computer and have another good go at it ...

Strangely 'tasklist /svc' reveals that it is Process ID 708 that is sending smtp packets out.  ID 708 is Services.exe.  Tasklist reveals that Services.exe is a host for Eventlog and PlugPlay.

Sysinternals' Process Explorer with a filter of 'Path contains :smtp' also reveals ID 708.   But it also shows that the operation is TCP Disconnect and TCP Reconnect - which means the router block is working.   There's one inetinfo.exe in there sending smtp to localhost ... which throws the possibility of IISADMIN, SMTPSVC and W3SVC involvement.

Sysinternals' RootkitRevealer (RKR) has found quite a few results...
HKU\1-5-21-nnn..nnnn\UserAssist\{75048700-EF1F-11D0-9888-006097DEACF9}\Count\... ...HRZR_EHACVQY:uggc://vzntrf.tbbtyr.pbz/ ... etc
Further investigation discounts this as suspicious...
1) {75048700-EF1F-11D0-9888-006097DEACF9} is the CLSID for ActiveDesktop - so all these entries are Operating System entries.
2) HRZR_EHACVQY:uggc://vzntrf.tbbtyr.pbz/ is ROT13 encoding for UEME_RUNPIDL:http://images.google.com/

Also see Didier Stevens UserAssist Tool for an easier decryption of these UserAssist entries.

Null containing registry entries found by RKR such as in the hklm\security hive SAC* SAI* and SCM* have the same date (maybe) as the computer was installed... something to do with password hiding perhaps...

HKLM\Software .. Microsoft SQL Server .. apparently this is logged by RKR because SQL often changes quickly whilst the registry is being read.. so comparing it can often show innocent discrepancies.
HKLM\System\ControlSet1\Services\eyviuy
HKLM\System\ControlSet2\Services\eyviuy
These have the same names as the driver file (.sys) I found in the system32\drivers folder...

AVG has popped up .. The Trojans, Hiloti.AL Hiloti.AM have infected two files in System Volume Information\_restore folder...

I have given in and taken it back to a previous system restore point .. early last week.  If it gets re-infected then I'll have to ask someone to go and run Recovery Console and delete the offending system driver file and registry entry without Windows running.  

Saturday, 17 April 2010

Dell Server PE1800 2GB RAM - Windows SBS Server 2003 R2: One Hour to Reboot

This server has quite a few issues ... the main issue is that the Backup Exec 12 backups are failing each night...

This blog isn't meant to be an anti-AVG blog... but AVG on this client's server is again the wrong version.  There should be safety blocks, detect Exchange running or something in the installer of the normal version to prevent it being installed... I've no idea why it's on there... perhaps it was an upgrade path at one time...

Way to check on an Exchange server - (which took me ages to finally discover but minutes to check now) - Run Poolmon.exe -b from the Server Support Tools for Windows Server 2003... look for the tag, AvgU in the 'NonP' (non-paged memory pool)... it shouldn't be there on an Exchange server.

So now I've uninstalled AVG and am installing the correct Exchange server version.

Rebooting was interesting...  If only for this precise jump in time:


Event Source: EventLog
Event ID: 6006
Date: 17/04/2010      Time: 05:01:12
Computer: SERVER1
Description:       The Event log service was stopped.
--------------------------------------------------------
Event Source: EventLog
Event ID: 6009
Date: 17/04/2010      Time: 06:00:07
Computer:         SERVER1
Description:       Microsoft (R) Windows (R) 5.02. 3790 Service Pack 2 Multiprocessor Free.
--------------------------------------------------------
Event Source: EventLog
Event ID: 6005
Date: 17/04/2010      Time: 06:00:07
Computer: SERVER1
Description:     The Event log service was started.

Anyone seen anything weird like that?  Precisely one hour between Event Log stopping and starting?  Could it be some settting in the Dell ACPI implementation for cooling the fans in the server?  I'm working remote so there's no way to check that without getting my colleague to look on his next visit.

Perhaps so few people notice this because they are sitting next to the machine and get impatient after 15 minutes and hard reset it ... I sat watching Task Manager ordered by Processor Percentage, as you do sometimes (yawn), and noticed dsm_sa_datamgr32.exe popping up into second place every now and again.

That kind of activity is going to play cruel with Virtual Shadow Copy during Adv. File Option activated backups... perhaps why the backup sits either in Snapshot Pre-processing or on-Queue all night and all day... there's not a quiet period during which it can start taking a snapshot... PERC overload...

Services - DSM SA Data Manager: "Provides a common interface and object model to access management information about the operating system, devices, enclosures and management devices. If this service is stopped, several management features will not function properly."

Version installed: 5.8.0.4938 - 3rd Dec 2007... that's a bit old, me reckons... There most certainly should be an update... have to upgrade to Dell Open Manage 6.2 ...

In the mean time I've been warned by AVG 9.0 that the version of Roxio Installed on this server is out of date and might cause problems with AVG 9.0 and other programs... so I found a fix here (tho this might disappear soon) - it asked for reboot and I don't want to until I've installed AVG 9.0 so I'm going to have to ignore the warning from AVG 9.0... and install it .. I won't know if that link solved the problem with Roxio.  Best thing would be to remove Roxio 5... and install.. what is it now? 6, 7, 8, 9, 2010?   I don't know if they are using it or not, so I can't just remove it ..

Strange .. I have found OpenManage 6.1.0 files ... but 5.8 is installed... 6.1.0 files were from 21st Oct 2009... but not installed?  Was a disk space issue apparently...  After doing a WSUS Cleanup there's 6GB free on F: ... so should be ok to install this time...

Seems like Dell brought out an update to Open Manage in Dec 2009... 6.2.0 .. fixes a leak or crash in DSM SA Data Manager ... so I'll update that...

Server updated and now rebooting...  Will it take another hour between the EventLog Stop and Start events in the Application Log?  Let's see...

Monday, 29 March 2010

iPhone and Exchange ActiveSync: Incorrect AVG Version

Well, having discovered the server was losing non-paged pool resources I used poolmon and watched the resources over a couple of days...

AvgU was rising along with File and Irp.

If you open a cmd window and cd to c:\Windows\System32\drivers ... you can run
'findstr /m /l AvgU *.sys'
That should return just the file name of any sys file that contains the literal string 'AvgU'. It returned 'Avgtdix.sys'.

Further googling... Avgtdix.sys is a Network Connection Watcher. This file should NOT be installed on an Exchange Mail Server.

I have uninstalled AVG from the server and installed the Exchange Mail Server version of AVG (paid).

No Avg tags are showing up in the Non-Paged Pool... there's no Online Shield or Email Scanner in this version...

I will have to monitor the server for another few days...

Friday, 26 March 2010

iPhone and Exchange ActiveSync: not solved - reoccurred

So I got a call that the iPhone stopped getting emails... I didn't check till they went home in case I needed to restart their server out-of-hours ...

When I did connect I decided to run MS Exchange Best Practices Analyser... (search on ExBPA).

I knew one of the Best Practices was having an Application Log size in Event Viewer of 40Mb - so I decided to alter that myself first whilst I was waiting... I must have dozed off at that point ... suddenly I was disconnected... RealVNC was suggesting I was connected - it wasn't asking for my username and password - but my connection was getting refused by the server - 'Read/state - Connection disconnected by peer (10054)'

Somewhere I read that 10054 'read/state' might mean that the Application log size was too small and so the server could not accept connections... that was close? Did I type 40kb? I couldn't connect via VNC... but I couldn't connect either via VPN... could that cause that too?

I had to wait until the morning and catch the first person into the office... when I called they said they couldn't log in either... they could log into the server terminal... a quick look at the System log revealed srv errors related to non-paged pool memory.... I got them to reboot so we could all log in ...

When non-paged memory has only 20Mb left then Windows shuts all connections down... IIS6, HTTP.SYS, users are logged out, VPN connections shut down, RealVNC connections... and so on... it does that to prevent resources becoming so low that the system crashes... you are forced to log into the server and sort it out.

David Wang's HowTo post helped point me to Poolmon.exe - a tool that monitors which components are using Paged and Non-Paged memory... Poolmon.exe is one of the Windows Support Tools and can be downloaded from Microsoft.

Once installed, open a cmd window and type: poolmon -b to list by bytes - which is the column to watch...

Because my client's server has just booted this will just tell me initial values... so I am saving to a file: poolmon -b -n datetime.txt. Then I'll import the file into Excel and time stamp each row. I'll run poolmon every now and again and see which values are changing and which aren't.

So far I've got several culprits:
  • Irp - - Io, IRP packets
  • File - - File objects
  • AvgU - an AVG component
  • Ntfr - ntfs.sys - ERESOURCE (not to be confused with NtfR)
  • MmCa - nt!mm - Mm control areas for mapped files
the largest riser is AvgU, followed by File. That's not surprising considering everyone had been logging in and were now accessing their files...

An hour later:
  • FMsl - - fltmgr.sys - STREAM_LIST_CTRL structure
  • File - - File objects
  • AvgA - an AVG component
  • AvgU - an AVG component
  • Ntfr - ntfs.sys - ERESOURCE
  • MmCa - nt!mm - Mm control areas for mapped files
My bet is that AVG is going to be the culprit... I will stop working on this now. Later I will take one more reading, but I may just remove AVG and install the correct server version - this version looks different to the other servers that I get to look at.

Saturday, 20 March 2010

Exchange: 13 month old email received on Blackberry

A client keeps on receiving batches of 10 or so emails on their Blackberry via Exchange... those forwarded to me (on the 19th March 2010) were dated 16th Feb 2009 and 19th October 2009.

The client has Microsoft Exchange 2003, AVG for Exchange, MAC Entourage, Blackberrys using Blackberry Internet Service (BIS).

I have connected up to Outlook Web Access and checked that the emails received are still in the user's Inbox... so there is some kind of periodic sweep over a batch of emails in the Mailbox every now and again...

First look... is AVG again... perhaps biased by not looking at AVG in the last problem (with iisadmin and https services not restarting)... version is 9.0.272... all fairly up-to-date...

Email Scanner for Exchange settings VSAPI (in Advanced Settings -> Server Components) selected components are:
  • Background Scan - On
  • Proactive Scan - Off
  • Scan RTF - On
  • Number of Scanning Threads: 9
  • Scan Timeout: 180
A background scan could sweep through email messages... I need to find out what AVG says this does...
Background scanning is one of the features of the VSAPI application interface. It provides threaded scanning of the Exchange Messaging Databases. Whenever an item that has not been scanned before is encountered in the users mailbox folders, it is submitted to E-mail Scanner for MS Exchange to be scanned. Scanning and searching for the not examined objects runs in parallel. Note: A specific low priority thread is used for each database, which guarantees other tasks (e.g. e-mail messages storage in the Microsoft Exchange database) are always carried out preferentially.
So... there's a background process, running on a low-priority thread, meaning it'll give up processor time to anything with a higher priority... the timeout is 3 minutes per email - (that's the maximum time the scanner can spend on any one email)... on a busy server that could take a long time to scan a few large emails (at the least 20 emails per hour (unless Exchange is busy sending and receiving emails or responding to lots of Entourage requests from around the Office?)...

So my theory at the moment is that the low priority thread that VSAPI background scanning is working on is having to give way to other higher priority threads...

Another factor is that the client's office is all Mac .. all Entourage... and Entourage talks to Exchange differently than Outlook does... does this explain why other clients don't have a problem with their Blackberry synchronising? I guess I have to look into VSAPI a little...

This from the MS Exchange Team Blog is useful background (Parts 1,2,3):

First, this advises (in Part 2) switching on Medium Diagnostics Logging on Antivirus Scanning...

In Exchange Admin -> Servers -> Server -> Properties -> Diagnostic Logging tab -> in services: MSExchangeIS -> System -> in categories: Antivirus -> Set Medium or Maximum logging level -> Click OK to exit...

Wow... just noticed while in Exchange Admin ( -> First Storage Group -> Mailboxes) that this particular user's Mailbox is 10GB... the largest of all the users on that network, next largest is 6GB then 4GB... So that is now a potential factor in slow Background Scanning of this mailbox and old emails on her Blackberry...

Diagnostic logging on ... also the Exchange Blog mentions that when an email is scanned it is stamped - "At the completion of the scanning process ptagVirusScanningStamp is updated reflecting the results of the scan. This property holds information such as the vendor, version, scan results, and miscellaneous information regarding the last scan of the item" - Perhaps if the email has never been scanned, this stamp affects the sync?

It's been a few minutes now... so time to check Event Viewer for Antivirus events... (darn.. I reckon I need to restart MSExchangeIS service!)

These came in long before I switched on Max logging level:

Event Type: Error
Event Source: MSExchangeIS
Event Category: Virus Scanning
Event ID: 9581
Date: 20/03/2010
Time: 12:44:23
User: N/A
Computer: SERVER
Description:
Error code -536768764 returned from virus scanner initialization routine. Virus scanner was not loaded.

So if the scanner has been unable to initialise on a regular basis... I should take a note of the dates and times of all Events with ID 9581 - they might tally with old emails being sent - or they might not... I find it's good to have some solid dates and times... It might also happen at a particular time of the day or periodically - so making a note of times and dates helps to see a pattern if one exists... Or simply filter all Events in the Application Log by ID 9581...

ok... these messages occur at 00:44am and 12:43pm every day for the last 3 days 17-20th March 2010; then on the 5th Feb 2010; 16th to 22nd Dec 2009; and 9th/10th Dec 2009

I've got to presume that the rest of the time MSExchangeIS antivirus scanning was working fine... On the days it doesn't work, something regular is interfering... like backups? ... but there are also other times... on 4:44am on 18th Dec... and two this morning at 5am and 8am... which is perhaps when I ran the Exchange Best Practice Analyzer... or Trace Analyzer...

Microsoft Best Practice Analyzers web site...

.. what is happening at 12:43 and 00:44? First check backup times, particularly the time that the Exchange backup begins...

It appears that Backup Exec (12) is half way through backing up (with GRT enabled) at 00:44... but it starts an hour earlier and finishes afterwards... and it can't explain the 12:43pm failure...

Ah... AVG Antivirus Update Manager is set to check for virus updates every four hours and the last one was at 12:44pm ... That explains the 4:44am one on the 18th Dec, maybe the 8:42am this morning... but not the 5:09am... perhaps it was when I was playing around with AVG Email settings... ? possibly...

Is event 9581 in 'MSExchangeIS\Virus Scanning' relevant? Time to ask the client for as many dates and times that they received these old emails on their Blackberry... if they can... See if the dates tally with the dates of these errors at all... and also find out whether any of the old emails received were dated after the 9th December... (perhaps they used Blackberry Enterprise Server (BES) before and Internet Service (BIS) after then...?)

Whilst digging around on the Exchange Team Blog... I discovered a post about a change made in 2006 to Exchange that could affect Blackberrys and other services - connected to 'Send As' permissions... I'm not sure that applies here but for anyone with that problem here...

they point to kb article - 912918 - Users cannot send emails from a mobile device or from a shared mailbox in Exchange 2000 or 2003 - it mentions that it would break sending emails if you use BES... (perhaps that accounts for an error I saw back in 2008...)

And - Send As permission behaviour change in Exchange 2003
A fix has been released that changes the behavior of the "Full Mailbox Access" feature in Microsoft Exchange Server 2003. Prior to this change, any user with the “Full Mailbox Access” permission for a mailbox also had the ability to “Send As” the mailbox owner.
... the script is not straightforward.. you have to read the first kb article to know how to use it properly... there's no '-?' switch... - doesn't appear to do anything on our server.. no output...

I'm going to leave the server running for a bit with MSExchangeIS Antivirus logging enabled... Seems that Antivirus starts scanning the Mailbox and Public Folders at 00:44 - Public Folders is over in 6 minutes and Mailbox goes on for between 4 and 19 hours. It then starts again... but is quicker. Could be that it only does a full scan if the virus signatures are updated in AVG - then it has to re-scan everything that it has marked suspect.

Going to leave this till I hear back from the user... with diagnostics running on Virus Scanning I'll be able to see a connection next time... or not...

Friday, 19 March 2010

Exchange ActiveSync and the iPhone: Problems Connecting (solved)

I think I have just about tried every trick in the book to get these iPhones connected to my client's SBS 2003 server... from deleting the Virtual Servers in Exchange Admin, removing them from the IIS MetaBase, DS2MB and letting SBS recreate them, checking through all the authentication and access settings, testing SSL, check User-level ActiveSync settings/Outlook Mobile Access and Global-level ActiveSync/OMA settings...

Up for a breather... and ... every time I ran IISRESET it would hang on HTTP SSL service - HTTP Filter.. so then there's no way that IISADMIN would restart... and then IIS is down because W3SVC won't come back up... so it meant a reboot every time...

[One of the problems on the SBS was that 2 updates had failed... .NET 2.0 SP2 (KB976569) was failing over and over - it appeared to take hours (14 hours - I tried to cancel the update and reboot the server - I should have made sure the Update was stopped before rebooting - since I was working on a server 200 miles away I had to wait till the client came in to find the server stuck on shutdown before I could get access again.) - after some investigation into the log files I found it failed because IISADMIN did not restart when asked. The other update was an Intelligent Message Filter for Exchange 2003 SP2 update from Feb 2010 (KB907747). Neither told me why they failed. Both failed due to IISADMIN not restarting (due to HTTP SSL not stopping)... I had to disable IISADMIN, reboot, install the updates, re-enable IISADMIN, reboot. And try to figure out why HTTP SSL was not stopping in a timely fashion.]

So my main focus became digging around in IIS... making sure the OWA and OMA applications in Default Website were attached to the correct ApplicationPool, ExchangeApplicationPool or the ExchangeMobile one...

Errors 3005 3007 for ActiveSync... really just tell me something's up... maybe a timeout... maybe the server is overloaded... but for a big server like that with a handful of users ... hardly likely... must be a setting somewhere... Number of concurrent connections .. in Performance on Default Website properties... that did something.

Check the log files under C:\Windows\System32\LogFiles\HTTPerr and ..\W3SVC1 - go to the end of each file and look for PROPFIND and POST and GET statements from ActiveSync... go to the end of each line and check if you get 409, 207, 403 ... if you're not getting a nice round number like 200, 400 etc then something's up... and it's another pointer in some direction...

You can view which devices have connected by connecting and visiting (using any browser):
https://mail.mydomain.com/exchange//NON_IPM_SUBTREE/Microsoft-Server-ActiveSync

Having fixed a few glitches here and there... and having been up all night and day... mostly waiting for reboots... something was holding HTTP open... and slowing this process down.. people don't reboot every time they make a change to IIS? duh....

I started to run combinations of the installed programs through Google... and there .. lo and behold ...

Antivirus... AVG ... You should not install AVG Online Shield, AVG Firewall and Email Scanner on a Windows Server running Exchange (and definitely not Exchange ActiveSync)

The Online Shield scans HTTP, HTTPS traffic - it has a hook into the HTTP SSL/HTTP Filter service ... and was not specified as a DependsOnService (or vice versa...) so there's no call to it to stop and start... HTTP SSL sits and waits for AVG Online Shield to stop using it for AVG to stop ... but that won't happen...

Avg Email Scanner just adds another layer around POP3 and SMTP ... with it scanning ports 25 and 110 on a machine with Exchange running... adds another layer in the timeout values possibly...

Once Online Shield and Email Scanner were switched off ... IISRESET worked without rebooting .. what a joy I did it several times over and over ...

An hour or two later the client texted me to say his iPhone was getting emails ... marvellous...

These websites were helpful:
Henrik Walther's in-depth Chapter 5 from his book "Securing Exchange Server 2003 & Outlook Web Access" - perfect for understanding the nitty-gritty bits of Exchange HTTP Virtual Folders etc...

AVG - What AVG components are not designed for server operating systems?

Microsoft's Exchange ActiveSync Test Website:

Microsoft Exchange ActiveSync Administration Tool ... a tool that you can use to delete old devices from ActiveSync Administration Tool... I should let my client know about it since his old iPhone that he replaced in December is still in the system.
Microsoft Exchange ActiveSync Certificate-Based Authentication Tool:

GoDaddy SSL Certificates - works well with most PDAs, iPhone, Windows Mobile, Android and Exchange ActiveSync - also quite cheap at the moment for secure authentication...

If anyone still needs help ... leave a comment...