Sunday, January 20, 2008

AMD/Intel Q4 2007 Earnings – Where Are We?

AMD's recent performance has been a study in disappointment. Earlier, SSE 5 and microbuffered memory had looked so promising for 2009. Now, it appears that these may not arrive even in 2010. Consecutive quarterly losses, low clocks, and the TLB bug have added to AMD's pain. In contrast Intel has rebuilt its revenues, upgraded C2D to the G0 stepping and delivered the 4-way capable Caneland platform. Intel has had some difficulty with 45nm Penryn but this hasn't mattered much with the G0 65nm quad cores available.

Sometimes things are a bit counter-intuitive. For example, I've seen many assume that AMD's fortunes began to boom in 2003 with the introduction of K8. In reality, 2003 was worse for AMD than 2002. 2004 was better but AMD suffered from volume limitations due to the larger K8 die. AMD's volume share stayed at about 16% from 2002 all through 2004. It wasn't until 2005, fully two years after the introduction of K8, that AMD began making strides in the market. We see the same pattern repeated by Intel where 2006 was far worse than 2005 in spite of the introduction of C2D. To understand just how well Intel is doing we need to compare with 2005, not 2006. The numbers show that Intel is doing well but hasn't quite recovered to what it had in 2005:

2005:
Processor Earnings – $28.1 Billion
4th Quarter Cash and Short Term Investments - $12.8 Billion
Average Gross Margin – 59.3%

2007:
Processor Earnings – $25.9 Billion
4th Quarter Cash and Short Term Investments - $12.8 Billion
Average Gross Margin – 51.9%

2008 Outlook:
Predicted Average Gross Margin - 57%

Intel's Outlook makes it clear that they do not expect to return to the Gross Margins of 2005. In fact, 57% doesn't even match the 58% Average Gross Margin of 2004. That point is quite puzzling. The common reasoning has been that as Intel switches to 45nm their costs will drop dramatically. This should easily be enough to boost the Gross Margin up from the current 58% to above the 2005 levels. But, Intel says this isn't going to happen and Intel's prediction of an Average Gross Margin of 52% for 2007 was right on mark. Apparently, Intel either expects that ASP's will drop or costs will rise enough to counter the decrease in costs from 45nm. This again goes against common reasoning which assumes that Intel's higher clock speeds insulate them from pricing pressure from AMD. This further goes against Intel's own statements of reorganization which were supposed to result in another round of layoffs in 2008. I think we now have to assume that Intel's reorganization has halted halfway through and that nothing further will be done. This however disagrees sharply with the inclination by many to label Intel as “lean”. Intel cannot be lean with only half of the reorg in place; Intel still has some love handles.

If we look at the facts instead of simply talking from our gut reaction this does make sense. AMD had gained a lot of server share in 2006. However, Intel took most of this share back at the end of 2006 dropping AMD's share by 40%. AMD's server share average for 2007 has been only 14.2% compared to the 23.5% average for 2006. Intel has held onto this share due both to the value of C2D dual core and having a monopoly on quad core. Unfortunately, the problem with being on top is that you can only go down. And, Intel now faces substantial quad core server pressure from AMD. AMD moved some 130K Barcelona server chips in 2007 and this will only increase. Secondly, AMD has substantially undercut Intel in terms of lower clocked quad core offerings. These lower clocked quads are also much more immune to Intel's lower 45nm power draw which is more dramatic as we increase clock. According to AMD, the high volume range for servers is 2.3Ghz. If this is true then Intel is not at all immune to server pricing pressure and we know that servers are Intel's highest ASP. I would expect that AMD will move back up to its 2006 server share in 2008. We also know that AMD's Griffin/Puma mobile platform will be out soon and this should put pressure on Intel's mobile segment which is the second highest ASP. It then begins to make some sense when we realize that the segment where Intel will still dominate will be desktop which is the lowest ASP. Even so, AMD did report some increase in desktop ASP in Q4.

The talk about an AMD takeover or bankruptcy is never ending. A takeover is technically and financially possible except that I can't think of a single company that wants to get into the frontline processor business and slug it out with Intel. Motorola was the last challenger but they divested their processor division as Freescale in 2004. Some might point to VIA. Well, VIA's parent company could pull off the finances but they've never had a frontline processor. Although Cyrix was initially competitive in 1995, Cyrix's position had slipped a generation by the time they were acquired by VIA in 1999. Secondly, VIA never manufactured Cyrix processors; they instead chose the Centaur architecture from IDT which was never frontline. Presumably if VIA had wanted to be competitive they would have worked on a new generation of the Cyrix design which would have been released around 2002 or 2003. This never happened so VIA as buyer is very unlikely.

Some have suggested nVidia or Samsung as a buyer but neither of these companies has any experience with processors. Samsung does memory products which is a long way from the logic intensive design of cpu's. These don't overlap at all in terms of manufacturing or design methodology. NVidia does chipsets and graphics and nVidia would be an obvious monopoly conflict due to AMD's ATI holdings. No doubt someone will mention IBM. However, IBM could never take over AMD's markets since IBM's use of AMD server processors would make it a competitor with its own customers. This would mean giving up the most profitable market segment of AMD's line and would probably impair volume with customers like HP, Gateway, and Dell that also produce servers. In other words, IBM would lose its current R&D support from AMD (the largest of all its partners), lose the technical support that AMD provides to Chartered, and lose AMD support for APM. The loss of volume would also make AMD processors cost more. Thus, IBM would turn a supporting partner and Intel hedge into a money losing niche processor division. Economically, this makes no sense at all. So, an AMD purchase is probably somewhat less likely than, say, striking oil in your backyard.

AMD managed to increase its revenues every quarter in 2007 as well as its Gross Margins. AMD finally managed to trim the quarterly loss to $164 Million. Now, I have to be honest and mention that AMD's current value per share of stock is worse than it was at the end of 2003. Adjusting for the change in number of shares AMD was comfortably higher by about $1.2 Billion but the $1.6 Billion goodwill charge has now brought this about $400 Million under. This is cause for concern if AMD's loss increases again in Q1. Supposedly AMD will break even in Q2. AMD's outlook though is oddly different from Intel's. Whereas Intel predicts a stagnant Gross Margin at a lower value than even 2004, AMD predicts that the current level of 44% will hold during the first half of 2008 but increase to 50% in the second half of 2008. So, why would this be? First of all, by Q3 AMD should finally have at least 2.8Ghz chips. This would cover nearly the entire range since it looks like Intel won't go higher than 3.16Ghz. There is also the question of 45nm from AMD. Statements made by AMD during the Earnings Call include:

Ruiz
“look forward to being able to ramp 45 nanometer aggressively in the second half of this year.”

Meyer
“Followed on in the second half of the year by 45 nanometer.”

Meyer
“we’ve got internal samples of our 45 nanometer microprocessors, we’re putting them through their paces currently and we’re on track to, the plans we talked about in the past which is to start our ramp in the first half of this year and ship revenue product in the second half of this year.”

I have seen some attempt to spin these statements to mean that AMD will begin producing 45nm in the second half with actual delivery in 2009. This, however, matches neither what AMD said nor common sense. Someone might mistakenly take Ruiz's comment to mean initial production however this is completely countered by Meyer's statement. Secondly, AMD first produced Brisbane samples in Q2 2006 and delivered product in Q4. If AMD has samples in Q1 then product delivery should by Q3. So, the most reasonable assumption is that AMD will begin production in Q2 and deliver some small amount in Q3. This amount is unlikely to have any effect on revenues but the Q4 volume should be more significant. Likewise, I doubt Intel's volume of Nehalem in Q4 will have any effect on revenues. So, adding in AMD's mobile and higher clocked quad cores in Q3 along with some 45nm in Q4 I could see an increase in Gross Margin in the second half. Finally, this begins to make sense. Right now, AMD overlaps the highest volume range and should increase its overlap during 2008. This would put pricing pressure on Intel with most likely some loss of ASP and this would offset the 45nm savings. On other hand, AMD's ASP's are very low so it can only move up. It also makes sense because 50% Gross Margin would still be significantly below Intel's 57%. This would be consistent with lower ASP's for AMD and higher costs due to a lower ratio of 45nm production. Some might also have noticed that AMD has scaled back its expected 2008 volume from 100 million units to 80-90 million units. This is undoubtedly due to the scale back in the FAB 38 ramp. This in turn has undoubtedly reduced AMD's target of 30% share for 2008.

I guess the bottom line is that as long as AMD can reach break even in Q2 it should be able to start gaining asset value again. And, it looks like AMD will finally get back to the processor revenues that it had in 2006. Judging from the current unit volume though I would say that AMD intends to start gaining again, especially in servers and Intends to get back to what it had in 2006. The really surprising thing is a comparison with 2001. AMD's volume share had been about 16% but this increased to 20% in 2001. With the release of the Northwood P4 and AMD's delay in getting to 130nm, AMD's share tumbled in early 2002. AMD left 2002 with roughly same 16% volume share that it had in 2000. It stayed at this level all through 2004 then creeped up to 18% in 2005. AMD's volume gain in 2006 was much more dramatic rising to 23%. If the 2002 pattern had been followed (as many predicted) then AMD should have lost this share again in 2007. When AMD's volume share tumbled in Q1 2007 many of these armchair experts gave themselves big pats on the back. However, they couldn't have been more wrong. AMD's volume share promptly rebounded and AMD has held onto its 2006 average for the last three quarters. What these self styled experts failed to take into consideration was that when the crash occured in 2002 AMD had only been at its high point for two quarters whereas when the crash came in 2007 it followed five quarters of good volume. In other words, the two are opposites. The high in 2001 was temporary whereas the dip in early 2007 was temporary. According to IDC AMD lost 0.4% unit share and is now at 23.1%. This is about equal to AMD's 2006 average. Overall, I expect AMD to gain back its 2006 server share and go above its previous mobile share while holding onto and increasing its current desktop share. This would mean an increase in unit share from the current 23% to perhaps 26% in 2008. This seems doable but falls quite a bit short of AMD's previous target of 30%.

Intel in 2008 should finally pull ahead of its previous processor revenue high of 2005. It's Gross Margins may not reach what they were but they will still be quite good and I'm sure Intel will continue increasing its cash and doing stock buybacks. By any account this should be good performance and a reasonable rate of growth. Pricing pressure in the second half of 2008 as AMD moves up above 2.6Ghz should produce some good values for buyers. We can look forward to seeing how well K10 scales and whether or not 45nm Shanghai produces any change in speed or any reduction in power draw. We should also be getting reviews of Nehalem in Q4.

Monday, December 24, 2007

AMD K10 Tries To Grow Some Teeth

It has been a long, long wait for proper testing on K10. The previous hasty reviews generally showed K10 doing poorly with user applications and pretty well with server applications. Unfortunately, there never seemed to be enough information to determine if K10 were working properly or contained some serious design flaw. With the update review at Legit Reviews we now have at least some information.

It is clear that the initial teething problems with ATI have been taken care of. The new Spider platform looks to be first rate and a good win for AMD. However, this platform does need a good processor and AMD has been slow to respond. The Phenom is apparently shipping but only as 9500 (2.2Ghz) and 9600 (2.3Ghz). Things don't really get interesting until 9900 (2.6Ghz) but this is either not available until Q1 or may even be pushed back to Q2 as recently stated by Digitimes. Of course, the same source also says that Intel is delaying the Core 2 Quad Q9300, Q9450 and Q9550 until late Q1. The difference is that AMD has not commented while Intel has confirmed the delay. However, Charlie at The Inquirer says that Intel is still have teething problems of its own with its 45nm process. I have to say that this again brings up the question of how much FUD is in factory previews. We have Intel showing Penryns clocked to 3.33Ghz and AMD showing K10's clocked to 3.0Ghz with neither chip anywhere on the horizon. I guess in this smoke and mirrors contest Intel does seem to be closer to reality with a 3.2Ghz Penryn QX9770 scheduled maybe end of Q1 while AMD hasn't yet indicated when 2.6Ghz will be out. And, of course, with Intel's current lead the question of a delay in 45nm is moot unless AMD is planning to essentially skip 65nm versions and start producing 45nm in Q2 (delivery in Q3). That is about the only way I could see AMD getting back on track.

Okay, back to our original point. Assuming AMD manages to get real hardware out the door someday and make the 9900 more substantial than vaporware what do we have to look forward to? Is 9900 a contender or joke? This was the original Sandra XII scores showing very poor performance for K10 in both Integer and FP (the red bar). Such poor performance suggested a serious problem with K10. However, the new Sandra XII SP1 scores show double the previous scores removing any question of a serious design error in the memory section. The Sandra XII SP1 score of 10306 is fully 52% greater than Intel's QX9770. This explains why K10 does so well on memory limited benchmarks. Even more importantly however is the fact that it is 13% faster than 6400+ showing that K10's hardware changes to the memory controller do make a difference.

Next we need to revisit the two benchmark rumors that dogged K10 long before its release: Pov-Ray and Cinnebench. Early scores from these two benchmarks had Intel enthusiasts cackling and howling that K10 was no faster than K8. Neither was a proper benchmark but this didn't seem to stop the naysayers. We can see on the PovRay scores that 6400+ has to be 20% faster to get 2% better peformance than E6750. Thus C2D is about 18% faster than K8. However, in answer to the notion that Pov-Ray proves that K10 is no faster than K8 we see that there is no such proof. If we take the E6750 score and scale it to quad core at 2.4Ghz we can see that Q6600 scales 98% which is a very good score. However, K10 scales 102%. This suggest that K10 is at least 4% faster than K8. Since Q6600 shows no bandwidth limitations we know that the improvement is not due to memory. True, the change is small but for unoptimized code it is significant. The PovRay website also does not say what level of SSE is supported in the 64 bit version. Judging from the small increase from K8 to K10 and C2D it is clear that PovRay is not using 128 bit SSE operations. It does say that the 32 bit version only supports up to SSE2. It also does not say what compiler was used to create the executable. This too is significant since if Intel's compiler were used, K10 would take at least a 10% hit in speed. This could easily put K10 even or perhaps slightly ahead. Again, in terms of K10's speed relative to K8, the Pov-Ray Realtime scores are even more telling. Intel does manage perfect scaling with QX9770 versus E6750. However, K10 is fully 11% faster than 6400+ would be with perfect scaling. So, the Pov-Ray rumor is shelved.

Next we examine the contraversial Cinnebench scores. At the bottom we can see that K8 needs 20% more clock speed to run neck and neck with E6750. With this benchmark we see a problem though. K10 only scales 88% of 6400+. This isn't far off of Q6600's 90% scaling. However, when we look at the Penryn's we see something different. The 3.0 and 3.2Ghz Penryn's show perfect scaling from E6750. We know that this is not due to code tuning since Q6600 is slower. We know that it is not due to memory bandwidth since 9900 is slower. This pretty much only leaves cache size as the determining factor. Unfortunately, a benchmark that fits into Penryn's larger cache to run full speed is worthless for performance comparison. So, Cinnebench is still up in the air. However, given the improvements that we saw with Pov-Ray I think the Cinnebench scores can be ignored for now. When Shanghai is eventually released it will have a good deal more L3 cache than K10. If its Cinnebench scores perk up then we will know for certain that Cinnebench is useless due to cache tuning.

I could go over the rest of the benchmarks in the review but for the most part they are useless since they do not utilize all the cores. And, if we aren't using all the cores then why not just buy a dual core system? We can see that there are still questions of how benchmarks are compiled and that Intel may still be getting an edge due to the use of its compiler which is still heavily tilted in its favor. We can see that a lot of software still has not caught up to the use of 128 bit SSE codes which is the only way that C2D or K10 show their SSE strength. Older processors as far back as PIII work quite well with 64 bit SSE as does K8. And, there still remains the problem of multi-threading since a quad core is useless without quad threads. Nevertheless, this has answered a couple of questions. There is no big design defect in the K10 memory controller as earlier benchmarks suggested and K10 is clearly faster than K8 in terms of Pov-Ray. Unfortuately, we still have not turned up any real proof of whether K10's 128 bit SSE functions are on par with C2D's (and roughly twice as fast as K8's). Hopefully, this information will turn up eventually.

For now, all we can say is that assuming K10 is a signficant improvement on K8 AMD needs to get it out the door at speeds of 2.6 Ghz and above and sooner rather than later. For AMD's sake we also have to wonder if 45nm is truly still on track which should mean skipping some 65nm parts or whether this has slid as well. I suppose nothing but time will tell. Worst case for AMD would be Intel with 3.2Ghz and 3.33Ghz desktop Penryns and Nehalem launched to tackle the server market while AMD struggles to reach 2.8Ghz on 65nm with 45nm pushed back to 2009. Best case would probably be starting production of 45nm Shanghai parts in Q2 with delivery in Q3 with parts available in clock speeds of at least 3.0Ghz while still fitting within the normal TDP and Intel still at 3.2Ghz with Penryn. The sad part for AMD is that I could see either of these scenarios happening and neither one is particularly bad for Intel. We all know what AMD's New Year's resolution should be; we just don't know if AMD is capable of living up to it. C2D has some serious teeth and the fangs are just getting longer with Penryn; K10 with milk teeth is far less threatening.

Monday, November 26, 2007

AMD: All Dressed Up, But No Place To Go

AMD's current position is frustrating at best. Although AMD seems to have gotten its ATI offerings in order, its K10 offerings lag expectations in almost every way.

AMD's previous 65nm ATI 29xx GPU offerings looked a bit dismal compared to nVidia's. The high power draw and low performance justified more than one lukewarm review. The best ATI product, HD 2900XT, was only close to the nVidia 8800GT (at higher power draw) but nowhere near the 8800GTX. AMD's response to this seemed more than a little strange. They said that they intended to compete against GTX with Crossfire. However, it seemed difficult to imagine anyone really wanting double the power draw and heat of 2900XT. It seems though that with the new 38xx series on the 55nn process, AMD has gotten the power draw under control. And, with the reduced die size they can probably sell them at the reduced 2900 price and still make money. And, it looks like 38xx will actually make Crossfire a viable competitor against GTX.

AMD's new 55nm based 7xx chipsets look like winners of the first order. And, AMD's Overdrive utility looks very good. It is amazing to think about tweaking the overclock on a system with a factory control panel. The GPU's, the chipsets, and the Overdrive utility dress up any new AMD system fit for a ball. The problem is that these systems require an equally good processor but, so far, AMD has failed to deliver.

But, it isn't just one problem with K10; it is many. From lagging volumes to a clock ceiling of just 2.3Ghz to slow NorthBridge and L3 cache speeds to no clear indication of improvement over K8. Not only is there no clear indication of when things might be fixed there is not even a clear indication of what the problem actually is. Low volumes would suggest poor yields and ramping problems. Low clocks could suggest either process problems or architecture problems. The lack of 2.4Ghz speeds was blamed on a bug that won't be fixed until the next revision. This is in contrast with both reviews that overclocked to 2.6Ghz with no trouble and AMD's own 3.0Ghz demo. Under normal circumstances AMD should be able to deliver a demo chip in volume about 6 months later. That AMD does not appear ready to deliver 3.0Ghz K10's anytime in Q1 suggests a big problem of some kind. The poor performance also contrasts with statements from different sources who suggested that K10 had great performance at higher speeds.

The reviews aren't much help either. OC Workbench shows performance for Phenom close to that of Kentsfield at the same clock with a few exceptions. For example the Multimedia - Int x8 iSSE3 score indicates either a compiler problem or an architecture problem. Phenom's score is only about 1/3rd what it should be after turning reasonably good scores in the other categories. The Cinnebench scores are odd since Phenom appears to speed up more than 4X as all four cores are used compared to 3.53X for Kentsfield. A really good speedup would be 3.8-3.9X while 4.0 should be impossible. Something funky is definitely going on to get 4.35X. Phenom also falls off quite a bit in the TMPGenc 4.0 Express test.

The Anandtech Phenom Review suggests that Intel might be having some small difficulty with clocks on Penryn. However, this difficulty would be so small compared to AMD's that it is hardly noticeable. It only means that 3.2Ghz desktop chips won't be out until Q1. Since this is also when AMD is releasing 2.4Ghz this could easily put Intel at a 33% faster clock. Other statements in this review suggest that Intel may have streamlined its internal organization considerably which would also be bad news for AMD. This review also mentions that K10's L3/NB can't clock higher than 2.0Ghz. Anandtech does however confirm how nice AMD's Overdrive Utility is. The Anandtech benchmarks show generally slower performance for Phenom at the same clock as Kentsfield with higher power draw for Phenom. So, the Anandtech scores don't really show us where a problem might be.

I was hoping that when the November 2007 Top 500 HPC list came out that I would get some new test scores that would settle the question about whether K10 was better than K8. I downloaded the latest Top 500 list as a spreadsheet. I was delighted to see a genuine Barcelona score. HPC systems typically have well tuned hardware, operating systems, and software so benchmarks from these should be accurate. I was hoping the HPC numbers would avoid any question of unfavorable hardware, OS, or compiler.

I began crunching numbers. I normalized the scores based on the number of processors and clock speed. I was happy to see very tight clustering at 2.0 for dual core Opteron. Then I ran the numbers for Barcelona and it showed 4.0. This was double the value for dual core. This seemed to be a big problem to me since with twice as many cores and double the SSE width it would seem that K10's top SSE speed should be twice per core or four times larger for Barcelona. In other words, I was expecting something around 8.0 but didn't see that. This would suggest a problem.

However, I then ran the Woodcrest and Clovertown numbers. Woodcrest clustered tightly at 4.0. This was not a surprise since it too has twice the SSE width of K8. Unfortunately, the Clovertown numbers continued to cluster at 4.0. This was a surprise since (with twice as many cores) I was expecting 8.0 for Clovertown as well. So, unfortunately, these scores are inconclusive. K10 is showing the same normalized score as Clovertown but neither is showing any increase in speed over Woodcest. The questions still remain unanswered.

The bottom line is that AMD is having problems. Whether these problems are due to process, design flaws, or unfixed bugs is hard to say. AMD is also under the gun on time. For example if Phenom is only hitting 2.4Ghz in Q1 then that leaves no time to start production of Shanghai in Q2 for a Q3 release. I suppose it is possible that AMD could fix a serious design flaw with the Shanghai release but it is by no means certain. What is certain though is that AMD will have to do much better to have any chance of increasing its share up to profitable levels in 2008.