Showing posts with label Penryn. Show all posts
Showing posts with label Penryn. Show all posts

Saturday, January 17, 2009

Seeking A New System

Finally, information is arriving on Phenom II (Shanghai) versus Penryn and i7 (Nehalem). It's time to start thinking about a new desktop.

Generally when I buy a new system I end up getting something 3X more powerful than my old system. However, the last time I bought a new computer I wanted a notebook so I ended up getting a desktop replacement system with a 2.0Ghz mobile Athlon 64 which is only about 67% faster than my old P4 1.8Ghz system. This has put me a little behind so I would really like something about 9X faster than the old P4 to catch up. This requirement would be met with a tri-core of 3.0Ghz or a quad of only 2.3Ghz. There are no X3's available at 3.0Ghz yet but both Intel and AMD easily fulfill the quad requirement. However, I was reminded of Anandtech's past views about processor value after some back and forth with Johan De Gelas.

AMD's dual core Opteron & Athlon 64 X2 - Server/Desktop Performance Preview

"The real problem is that AMD has nothing cheaper than $530 that is available in dual core, and this is where Intel wins out. With dual core Pentium D CPUs starting at $241, Intel will be able to bring extremely solid multitasking performance to much lower price points than AMD"

Athlon Dual Core: Overclocking the 4200+

"The pricing, however, was a little hard to swallow with the range from just over $500 for the lowest-priced 4200+ to around $1000 for the top-line . . .

With the 4400+ sporting 1MB cache on each core, and only a few dollars more than the 512KB 4200+, we would suspect the 4400+ may well be the Dual-Core to buy
"

Affordable Dual Core from AMD: Athlon 64 X2 3800+:

"and today, AMD is launching it - the $354 Athlon 64 X2 3800+; the first somewhat affordable dual core CPU from AMD.

The cheapest dual core Pentium D processor could be had for under $300, yet AMD's cheapest started at $537. Intel was effectively moving the market to dual core, while AMD was only catering to the wealthiest budgets.

The Pentium D 820, running at 2.8GHz and priced at $280, offered the most impressive value that we've seen in a processor in quite some time
"


Of course, Anandtech has been leaning towards Intel for so long, I doubt they even remember what they said back in 2005. I suppose this could be why Anandtech never complains about price with Intel's Nehalem or Q9650 as they did (over and over) with AMD's X2. So, a $500 pricetag is "catering to wealthiest budgets", $354 is only "somewhat affordable", $280 is "impressive value", and more L2 cache is better at $44 because this is "only a few dollars more". I suppose someone could argue that inflation has raised the price point since 2005 however this has been more than offset by the decrease in system prices. So, let's see what happens if we use Anandtech's past, vigorously defended (but now apparently forgotten) criteria with today's chips:

We immediately eliminate the Intel i7 940 at $565 and Q9650 at $550 plus anything higher. At first glance, i7 920 would seem to be okay at $295. However, by the time we allow for the extra money required for the more expensive X58 motherboard and DDR3 memory we are again back up to $500, so it gets eliminated as well. This leaves Q9550 and below. However, Anandtech also preferred the 4400+ because it had twice as much L2 cache as the 4200+. This would make it very difficult to choose the $250 2.66Ghz Q9400 with 6MB L2 over the faster $295 2.83Ghz Q9550 with 12MB L2. In fact, the difference in price is almost identical to the difference between the 4400+ and 4200+. And, if double the cache is good at $44 then presumably triple the cache would be equally good at $66. But with only 1/3rd the cache the Q8300 is only $57 cheaper so it gets eliminated too. For Intel this only leaves the $190 2.4Ghz Q6600 with 8MB L2 and the $180 2.33Ghz Q8200 with 4MB L2. So, we have a choice between the outdated B3 stepping Q6600 and the newer 45nm Q8200 with half as much cache.

1/20/2009: The prices changed since I wrote this so I'll amend my comments. The Intel 3.0Ghz Q9650 is now "somewhat affordable" at $350and therefore at the top of the midrange. It is not a great bargain since 20% more money only gets you 6% more performance. Or, at least it does as long as the code can run from the L2; Q9650 shares the same FSB bottleneck as Q9550. The price of Q9550 has hardly change and at $290 is just $5 cheaper than it was. The Q9400 still gets eliminated because of the $50 price difference and half the cache but we retain Q8300 at $205 because it is now $85 cheaper than Q9550.

The problem is that AMD had price drops as well cutting Phenom II by $40. This makes Q9550 less of a bargain since PII 940 is now $55 cheaper. This is enough for a nice upgrade on the video card and to be honest I'd rather have a PII 940 with an HD 4850 than a Q9550 with an HD 4830. The drop in price with PII 920 to just $195 now makes it quite a value and throws a big monkey wrench into Intel's prices. Q9400, Q8300, Q6600, and Q8200 are all now about $45 too expensive. A choice between a cache crippled 2.33Ghz Q8200 for $170 and a PII 920 at 2.8Ghz for just $25 more is a no brainer. On the other hand, a Toliman system would be $50 cheaper again allowing a nice video upgrade.


If I were looking for a stopgap system I might go with the Intel $190 Q6600 or $180 Q8200. The AMD Black Edition Phenoms for $150 these days are also a bargain as is the $119 2.4Ghz 8750 BE Toliman tri-core when matched up with a 750 southbridge. I would be comfortable with any of these chips on a bargain or stopgap system. However, since AMD tripled the L3 on with its newer chips we are again pushed towards the $235 Phenom II X4 920 at 2.8Ghz or the $275 Phenom II X4 940 BE at 3.0Ghz for a midrange system.

I don't usually bother overclocking but this is a pretty simple operation if you match a 750 southbridge with the AMD quads. If I were comparing with the 9950 BE or 8750 BE in terms of overclocking I might give the B3 Q6600 the benefit of the doubt. But in the company of the newer 45nm chips the vintage Q6600 is clearly out of its league in terms of overclocking. Given a choice between the 920 and 940 AMD quads I would have to choose the 940 because its Black Edition unlocked multiplier would make it easier to overclock.

The Intel Q9550 is technically faster however it gets choked by its slower 1333Mhz FSB. To match the AMD 940 it would need 2133Mhz.However, even the still slower 1600Mhz FSB is only available on the insanely overpriced QX9770 for $1,460. Intel integrated graphics still lag behind both nVidia and AMD, so an nVidia motherboard is required for Intel unless I want to add a discrete graphics card from the start. The Dragon platform is the obvious choice for AMD although nVidia options are only a little worse in terms of integrated graphics.

For a discrete graphics card the 4000 series would be the obvious choice for something from AMD (ATI). I would probably be looking at a $140 HD 4850 or the $100 HD 4830 (which is about 75% of the 4850's performance). The nVidia equivalent of 4830 would be the 9800GT at $115 and the 4850 equivalent would be the 9800GTX+ at $165. The problem is that for $165 I can get a 1GB 4850 instead of a 512MB. And, if I move up to the 1GB size of 9800GTX+ I'm at $185 which is only $5 less than the much more powerful HD 4870. The GT260 is more expensive for less performance than the 4870 so it isn't worth considering.

I also wasn't exactly surprised with Anandtech's conclusion in AMD Phenom II X4 940 & 920: A True Return to Competition:

"Compared to the Core 2 Quad Q9400, the Phenom II X4 940 is clearly the better pick. While it's not faster across the board, more often than not the 940 is equal to or faster than the Q9400. If Intel can drop the price of the Core 2 Quad Q9550 to the same price as the Phenom II X4 940 then the recommendation goes back to Intel. The Q9550 is generally faster than the 940"

This conclusion is a bit out of whack since their own data shows that 940 compares favorably with Q9550. I would also guess that 940 would compare even better if Anandtech stopped avoiding mixed testing (for the past two years).

At any rate, I tend to agree with the guidelines that Anandtech used back in 2005. Today, I could easily recommend a Core 2 Quad Q9550 system with Intel motherboard if someone were planning to add discrete graphics anyway or with an nVidia motherboard for its better integrated graphics. An equally good choice would be a Phenom II 940 system with either discrete graphics or integrated graphics on a AMD Dragon platform. An nVidia motherboard would be adequate with Phenom II but if an older Phenom or Toliman were chosen then an AMD motherboard with 750 southbridge would be preferred.

1/20/2009: With the current prices the Q9550 is still acceptable but a worse bargain than a PII 940 system. In other words, with equivalent price it would be nearly impossible to assemble a Q9550 system equal to the PII 940 system. The drop in the PII 920 price has also greatly reduced the value of the lower end Intel quads and I could no longer recommend them even for a bargain or stopgap system. If you have money to burn, the Q9650 is less value for the money but would be an upgrade over a PII 940 or Q9550 system.

Monday, July 14, 2008

Reviews And Fairness Or How To Make Intel Look Good

I've had people complain that I've been too tough on Anand but in all honesty Anandtech is not the only website playing fast and loose with reviews.

Anand has made a lot of mistakes lately that he has had to correct. But, aside from mistakes Anand clearly favors Intel. This is not hard to see. Go to the Anandtech home page and there on the left just below the word "Galleries" is a hot button to the Intel Resource Center. Go to IT Computing and the button is still there. Click on the CPU/Chipset tab at the top and not only is the button still there on the left but a new quick tab to the Intel Resource Center has been added under the section title right next to All CPU & Chipset Articles. Click the Motherboards tab and the button disappears but there is the quick tab under the section title. There are no buttons or quick tabs for AMD. In fact there are no quick tabs to any company except Intel. Clearly Intel enjoys a favored status at Anandtech.

What we've seen since C2D was released is a general shift in benchmarks that favor Intel. In other words instead of shooting at the same target we have reviewers looking to see where Intel's arrows strike and then painting a bullseye that includes as many of them as possible. For example, encryption used to be used as a benchmark but AMD did too well on this so it was dropped and replaced with something more favorable to Intel. There has been a similar process for several other benchmarks. Of course now it isn't just processors. Reviewers have carefully avoided comparing the performance of high end video cards on AMD and Intel processors. Reviews are typically done only on high end Intel quad cores. The claim is that this is for fairness but it also avoids showing any advantage that AMD might have due to HyperTransport. It is a subtle but definite difference where review sites avoid testing power draw and graphic performance with just Integrated Graphics where AMD would have an advantage. They then test performance with only a mid range graphic card to avoid any FSB congestion which again might give AMD an advantage. Then high end graphics cards are tested on Intel platforms only which avoids showing any problems that might be particular to Intel. We are also now hearing about speed tests being done with AMD's Cool and Quiet turned on which by itself is good for a 5% hit. I suppose reviewers could try to argue that this is a stock configuration but these are the same reviewers who tout overclocking performance. So, by shifting the configuration with each test they carefully avoid showing any of Intel's weaknesses. This is actually quite clever in terms of deception.

As you can imagine the most fervent supporters of this system are those like Anand Lal Shimpi who strongly prefer Intel over AMD. I had one such Intel fan insist that Intel will show the world just how great it was when Nehalem is released. However, I have a counter prediction. I'm going to predict that we will see another round of benchmark shuffling when Nehalem is released. And, I believe we will see a concerted effort to not only make Nehalem look good versus AMD's Shanghai but also to make Nehalem look good compared to Intel's current Penryn processor. It would be a disaster for reviewers to compare Nehalem and conclude that no one should buy it because Penryn is still faster . . . so that isn't going to happen.

An example is that since AMD uses separate memory areas for each processor it needs an OS and applications that work with NUMA. In the past reviewers have run OS's and benchmarks alike oblivious to whether they worked with NUMA or not. If anything seems to be overly slow they just chalk it up to AMD's smaller size, lack of money, fewer engineers, etc. Nehalem however also has separate memory areas and needs NUMA as well. I predict that these reviewers will suddenly become very sensitive to whether or not a given benchmark is NUMA compatible and will be quick to dismiss any benchmark that isn't. This may extend so far as to purposefully run NUMA on Penryn to reduce its performance. This would easily be explained away as a necessary shift while ignoring that it wasn't done for K8 or Barcelona. That would be explained away as well by saying that the market wasn't ready for it yet when K8 was launched. That was what happened with 64 bit code which was mostly ignored. However, if Intel had made the shift to 64 bits reviewers would have fallen all over themselves to do 64 bit reviews and proclaim AMD as out of date just as they did every time Intel launched a new version of SSE.

We see this today with single threaded code. C2D and Penryn work great with single threaded code but have less of an advantage with multi-threaded code and no actual advantage with mixed code. It is a quirk of Intel's architecture that sharing is much more efficient when the same code is run multiple times. If you compared multi-tasking by running a different benchmark on each core Intel would lose its sharing advantage and have to deal with more L2 cache thrashing. Even though mixed code tests would be closer to what people actually do with processors reviewers avoid this type of testing like the plague. The last thing they want to do is have AMD to match Intel in performance under heavy load or worse still actually have AMD beat a higher clocked Penryn. But Nehalem uses HyperThreading to get its performance so I predict that reviewers will suddenly decide that single threaded code (as they prefer today) is old fashioned and out of date and not so important after all. They will decide that the market is now ready for it (because Intel needs it of course).

Cache tuning is another issue. P4EE had a large cache as did Conroe. C2D doubled the amount of cache that Yonah used and Penryns have even more. However, reviewers carefully avoid the question of whether or not processors benefit from cache. This is because benchmark improvements due to cache tend to be paper improvements that don't show up on real application code. So, it is best to avoid comparing processors of different cache sizes to see benchmarks are getting artificial boosts from cache. I did have one person try to defend this by claiming that programmers would of course write code to match the cache size. That might sound good to the average person but I've been a programmer for more than 25 years. Try and guess what would happen on a real system if you ran several applications that were all tuned to use the whole cache. Disastrous is the word that comes to mind. But you can avoid this on paper by never doing mixed testing. A more realistic test for a quad core processor is to run something like Folding@Home on one core and graphic encoding on another while using the remaining two to run the operating system and perhaps a game. Since the tests have to be repeatable you can't run Folding@Home as a benchmark but that isn't a problem since it is the type processing that needs to be simulated rather than the specific code. For example you could probably run two different levels of Prime95 tests in the background while running a game benchmark on the other two cores to have repeatable results. And, if you do run a game benchmark on all four cores then for heavens sake use a high graphic card like a 9800X2 instead of an outdated 8800.

Cache will be an issue for Nehalem because it not only has less than Penryn but it has less than Shanghai as well. It also loses most of its fast L2 in favor of much slower L3. My guess is that if any benchmarks are faster on Penryn due to unrealistic cache tuning this will be quickly dropped. That reviews shift with the Intel winds is not hard to see. Toms Hardware Guide went out of its way to "prove" that AMD's higher memory bandwidth wasn't an advantage and that Kentsfield's four cores were not bottlenecked by memory. But now that Nehalem has 3 memory channels the advantage of more memory bandwidth is mentioned in every preview. We'll get the same thing when Intel's Quick Path is compared with AMD's HyperTransport. Reviewers will be quick to point to raw bandwidth and claim that Intel has much more. They of course will never mention that the standard was derated from the last generation of PCI-e and that in practice you won't get more bandwidth than you would with HyperTransport 3.0.

I could be wrong; maybe review sites won't shift benchmarks when Nehalem appears. Maybe they will stop giving Intel an advantage. I won't hold my breath though.

Addition:

We can see where Ars Technica discovers PCMark 2005 error. Strangely the memory score gets faster when PCMark thinks the processor is an Intel than when it thinks it is a an AMD. Clearly the bias is entirely within the software since the processor is the same in all three tests:



I've had Intel fans claim that it doesn't matter if Anandtech cheats in Intel's favor because X-BitLabs cheats in AMD's favor. Yet, here is the X-Bit New Wolfdale Processor Stepping: Core 2 Duo E8600 Review from July 28, 2008. In the overclocking section it says:

At this voltage setting our CPU worked stably at 4.57GHz frequency. It passed a one-hour OCCT stability test as well as Prime95 in Small FFTs mode. During this stress-testing maximum processor temperature didn’t exceed 80ºC according to the readings from built-in thermal diodes.

This sounds good but the problem is that to properly test you have run two separate copies of Prime95 with core affinity set so that it runs on each core. This article doesn't really say that they did that. There is a second problem as well dealing with both stability and power draw testing:

We measured the system power consumption in three states. Besides our standard measurements at CPU’s default speeds in idle mode and with maximum CPU utilization created by Prime95

This is actually wrong; the Prime95 test they peformed was not maximum power draw; it was Prime95 in Small FFTs mode. But this doesn't agree with Prime95 itself which clearly states in the Torture Test options:

In-place large FFTs (maximum heat, power consumption, some RAM test)

So, the Intel processors were not tested properly. Using the small FFTs does not test either maximum power draw or temperature and therefore doesn't really test stability. If any cheating is taking place at X-Bit, it seems to be in Intel's favor.

Friday, February 29, 2008

2008: AMD Still Trailing

Intel is still moving along about the same as it has been making slow, incremental progress from the time that C2D was launched in 2006. It is clear that the increase in speed from Penryn doesn't match the early Intel hype but nevertheless any increase in speed is just that much more that it is faster than AMD. Likewise the tiny speed increases from 3.0 Ghz to 3.16 Ghz (quad core) and 3.4Ghz (dual core) are no doubt frustrating for Intel fans who would like more speed. On the other hand, AMD currently has nothing even close..

AMD will probably deliver 2.6 Ghz common chips in Q2. This chart at Computerbase claims AMD will release an FX chip in Q3. I'm not so sure about this because everything would suggest an FX of only 2.8 Ghz. This is probably the lowest clock that AMD could possibly get by with on the FX brand. There is no doubt that AMD needs a quad FX because people who bought FX in 2006 were promised upgrades and none have been forthcoming. Such a 2.8 Ghz FX would probably be clockable to 3.0 - 3.1 Ghz (with premium air cooling) based on what I've seen of the B3 stepping. This is probably the best AMD can do for now as I haven't seen anything that would suggest that B3 can deliver 2.8 Ghz as a common volume. This means that the poor man's version of FX, Black Edition will probably bump up to 2.6 Ghz as well. Intel seems to be somewhat behind in terms of 45nm but this hardly matters since their G0 stepping of 65nm works so well. But there is no doubt that AMD will be facing more 45nm Penryns in Q2. The shortages of chips have shielded AMD somewhat from increased presssure from Intel during Q4 and Q1 (although with Barcelona delayed server share may take another hit in Q1). However, as Q2 is the lowest volume of the year AMD will have to be aggressive to avoid a volume share drop during that quarter.

Probably, Fuad is closer to the truth of FX saying Q3:

The Deneb FX and Deneb cores, both 45nm quad-cores, are the first on the list. If they execute well we should see Deneb quad-core with shared L3 cache and 95W TDB in Q3. If not, the new hope will slip into Q4.

The timeline for FX being Q3 or maybe Q4 is not surprising at all. What is surprising is the idea that AMD's first new FX chip would be 45nm. If this is true then this would support the notion that AMD has suspended development of 65nm. But it would be surprising if 45nm could ramp that quickly.

The question then is what will happen in Q3 as AMD faces a steadily increasing volume of Penryn chips. The rumors suggest that AMD will not try to release a 65nm 2.8 Ghz Phenom. I'm not sure if this would then indicate that the 65nm process would hit a ceiling or whether this is to suggest that AMD will pursue these speeds with 45nm Shanghai. Another question is what 9850 might be. 9750 is supposed to be 2.4 Ghz while 9950 has been suggested to be 2.6 Ghz. So, would 9850 be 2.5Ghz perhaps? The topping out of the naming scheme does lend some credibility to the idea that AMD will suspend 65nm development and try to move to Shanghai as quickly as possible. Nevertheless, there is a big, big question of whether AMD could really deliver a 2.8 Ghz 45nm Shanghai in, say, Q3. Ever since the release of 130nm SOI, AMD's initial clock speeds on the new process have always been lower so there is a lot of doubt that AMD could reach 2.8Ghz on 45nm any sooner than Q4 2008. Nehalem will almost certainly be too small of volume in Q4 to be much of a factor. So, it looks like AMD's goal is to somehow get clock speed up and this seems even less likely with a mid year switch in process unless with 45nm AMD exceeds all past SOI efforts.

Early 2009 looks pretty good for Intel since it will not only have Penryn and Dunnington but increasing volumes of Nehalem. It still remains to be seen if Intel really will give up its lucrative chipset business on the desktop with Nehalem. It certainly seems that it wouldn't take much effort to modify an AMD HT based chipset to work with Intel's CSI interface. That would seem to remove a lot of Intel's current proprietary FSB advantage. On the other hand, with ATI out of the way this would seem to be the best time for Intel to face more competition in chipsets. Still, this does leave things a bit up in the air. If it becomes easier and cheaper to design chipsets as it surely would be if CSI is similar to HT then VIA might become more competitive. For AMD's part there seems little they can do in 2009 except try to ramp the clock speeds on Shanghai.

We have three other issues to talk about: one immediate and two longterm. The immediate issue is Hester's interview at Hexus where he mentions the slow clock speeds of K10. Basically, Hester says that the 65nm process is fine; it is a matter of adjusting some critical paths. I've seen this statement heckled by some who insist that you can't separate process from design. Curiously, these are the same people who also insisted that Intel's 90nm process was fine and that it was only a poor design with Prescott that was the problem. Anyway, this statement by Hester actually seems quite accurate to me. It was my impression that AMD had intended K10 to run at lower voltage which would have allowed higher clocks. This again seems to fit what we've seen with K10's limited by TDP. The reason for the higher voltage seems to be that the transistors don't quite switch fast enough and this causes some of the "critical paths" that Hester talked about to get out of synch. You could fix this at a low level by improving the transistors to get them back into spec with the design. Or, you could relax the timing on these critical paths which would get the design on spec with the transistors. Because 45nm is right around the corner it appears that AMD has decided to not expend more resources on 65nm improvement and will instead relax the timing. AMD's work on 45nm transistors will theoretically migrate down to 65nm, at least this is the theory of AMD's CTI (Continuous Transistor Improvement) program. However, we may now be entering a new era where improvements are so specialized that they may be unable to cross process boundaries as they used to and we may see AMD following Intel's lead. This would mean tighter design at the beginning of each process node and less reliance on later improvements.

The two long term issues concern the possibility of a New York FAB for AMD and the announcement on EUV. There are three questions about a NY FAB: Does AMD need it? Can they afford it? And, why NY instead of Dresden where FAB 30 and 36 are now? Need is most obvious because without a new FAB AMD's capacity will top out by mid to late 2010 unless the FAB 38 ramp is slower than expected. Affording is a big question but one that AMD can leave aside for now hoping that their cash situation will improve. The question of location is a curious one. One suggestion was that NY simply offered more incentives than Dresden but this by itself seems unlikely. In every case in the past Germany has shown itself more than willing to contribute money for AMD's FABs. So, the real reason for the NY location may have more to do with other factors. In fact, we even seemed to have some evidence of this from the EUV announcement.

"The AMD test chip first went through processing at AMD’s Fab 36 in Dresden, Germany, using 193 nm immersion lithography, the most advanced lithography tools in high volume production today. The test chip wafers were then shipped to IBM’s Research Facility at the College of Nanoscale Science and Engineering (CNSE) in Albany, New York where AMD, IBM and their partners used an ASML EUV lithography scanner installed in Albany through a partnership with ASML, IBM and CNSE, to pattern the first layer of metal interconnects between the transistors built in Germany."

Secondly, we need to remember that AMD only fell behind on process technology when it moved to 130nm in 2002. Prior to this AMD was doing pretty well. Although things seemed to improve after AMD's rocky transition to 130nm SOI AMD now seems to be falling behind again at 45nm. AMD used to operate its Submicron Development Center (SDC) in Sunnyvale, California. This facility was leading edge back in 1999. It surely is not lost on AMD that they have now surpassed IBM. Back in 2002 AMD only had a 200mm FAB while IBM had a more modern 300mm FAB as well as more capacity. AMD today has caught up in terms of FAB technology but passed IBM in terms of capacity. The big question for AMD has to be how badly IBM needs leading edge process technology and for how long. Robust server and mainframe chips need reliability more than top speed. Secondly, IBM has been steadily divesting hardware so one has to wonder when the processor division might become a target. Notice that in the above announcement the wafers had to be flown from FAB 36 in Dresden to New York. Given these facts I think it is possible that AMD wants to create another research facility at New York. I think this could serve both to tweak processes faster and optimize them better for AMD's needs as well as pick up any slack if research at IBM falls off. There has been no indication of this but it does seem plausible.

The recent EUV announcement is incomplete however. If we look at an IBM article on EUV in EETimes from February 23, 2007 we see that IBM very much wanted EUV for 22nm but figured that it wouldn't be ready in time for early development work.

The industry hopes EUV will make it into production sooner than latter, but the technology must reach certain milestones. ''I think the next 9 to 12 months are very critical to achieve this,'' said George Gomba, IBM distinguished engineer and director of lithography technology development at the company.

Twelve months from February 2007 would be now. So, what is missing from the EUV announcement is whether or not this recent test puts EUV on track for IBM for 22nm or whether it will have to wait for 16nm. A second question is why the test wafer was made at Dresden by AMD. If IBM had already tested its own wafers then why didn't it announce earlier? This could mean that AMD has decided to try to hit the 22nm node for EUV but that IBM has decided to wait until 16nm. If this is a more aggressive stance for AMD then it could mean that AMD will rely less on IBM for process technology for 22nm. This again would support the idea that AMD wants a new design center in NY. I think it is entirely plausible that AMD could surpass IBM to become the senior partner in process development over the next few years.

Sunday, January 20, 2008

AMD/Intel Q4 2007 Earnings – Where Are We?

AMD's recent performance has been a study in disappointment. Earlier, SSE 5 and microbuffered memory had looked so promising for 2009. Now, it appears that these may not arrive even in 2010. Consecutive quarterly losses, low clocks, and the TLB bug have added to AMD's pain. In contrast Intel has rebuilt its revenues, upgraded C2D to the G0 stepping and delivered the 4-way capable Caneland platform. Intel has had some difficulty with 45nm Penryn but this hasn't mattered much with the G0 65nm quad cores available.

Sometimes things are a bit counter-intuitive. For example, I've seen many assume that AMD's fortunes began to boom in 2003 with the introduction of K8. In reality, 2003 was worse for AMD than 2002. 2004 was better but AMD suffered from volume limitations due to the larger K8 die. AMD's volume share stayed at about 16% from 2002 all through 2004. It wasn't until 2005, fully two years after the introduction of K8, that AMD began making strides in the market. We see the same pattern repeated by Intel where 2006 was far worse than 2005 in spite of the introduction of C2D. To understand just how well Intel is doing we need to compare with 2005, not 2006. The numbers show that Intel is doing well but hasn't quite recovered to what it had in 2005:

2005:
Processor Earnings – $28.1 Billion
4th Quarter Cash and Short Term Investments - $12.8 Billion
Average Gross Margin – 59.3%

2007:
Processor Earnings – $25.9 Billion
4th Quarter Cash and Short Term Investments - $12.8 Billion
Average Gross Margin – 51.9%

2008 Outlook:
Predicted Average Gross Margin - 57%

Intel's Outlook makes it clear that they do not expect to return to the Gross Margins of 2005. In fact, 57% doesn't even match the 58% Average Gross Margin of 2004. That point is quite puzzling. The common reasoning has been that as Intel switches to 45nm their costs will drop dramatically. This should easily be enough to boost the Gross Margin up from the current 58% to above the 2005 levels. But, Intel says this isn't going to happen and Intel's prediction of an Average Gross Margin of 52% for 2007 was right on mark. Apparently, Intel either expects that ASP's will drop or costs will rise enough to counter the decrease in costs from 45nm. This again goes against common reasoning which assumes that Intel's higher clock speeds insulate them from pricing pressure from AMD. This further goes against Intel's own statements of reorganization which were supposed to result in another round of layoffs in 2008. I think we now have to assume that Intel's reorganization has halted halfway through and that nothing further will be done. This however disagrees sharply with the inclination by many to label Intel as “lean”. Intel cannot be lean with only half of the reorg in place; Intel still has some love handles.

If we look at the facts instead of simply talking from our gut reaction this does make sense. AMD had gained a lot of server share in 2006. However, Intel took most of this share back at the end of 2006 dropping AMD's share by 40%. AMD's server share average for 2007 has been only 14.2% compared to the 23.5% average for 2006. Intel has held onto this share due both to the value of C2D dual core and having a monopoly on quad core. Unfortunately, the problem with being on top is that you can only go down. And, Intel now faces substantial quad core server pressure from AMD. AMD moved some 130K Barcelona server chips in 2007 and this will only increase. Secondly, AMD has substantially undercut Intel in terms of lower clocked quad core offerings. These lower clocked quads are also much more immune to Intel's lower 45nm power draw which is more dramatic as we increase clock. According to AMD, the high volume range for servers is 2.3Ghz. If this is true then Intel is not at all immune to server pricing pressure and we know that servers are Intel's highest ASP. I would expect that AMD will move back up to its 2006 server share in 2008. We also know that AMD's Griffin/Puma mobile platform will be out soon and this should put pressure on Intel's mobile segment which is the second highest ASP. It then begins to make some sense when we realize that the segment where Intel will still dominate will be desktop which is the lowest ASP. Even so, AMD did report some increase in desktop ASP in Q4.

The talk about an AMD takeover or bankruptcy is never ending. A takeover is technically and financially possible except that I can't think of a single company that wants to get into the frontline processor business and slug it out with Intel. Motorola was the last challenger but they divested their processor division as Freescale in 2004. Some might point to VIA. Well, VIA's parent company could pull off the finances but they've never had a frontline processor. Although Cyrix was initially competitive in 1995, Cyrix's position had slipped a generation by the time they were acquired by VIA in 1999. Secondly, VIA never manufactured Cyrix processors; they instead chose the Centaur architecture from IDT which was never frontline. Presumably if VIA had wanted to be competitive they would have worked on a new generation of the Cyrix design which would have been released around 2002 or 2003. This never happened so VIA as buyer is very unlikely.

Some have suggested nVidia or Samsung as a buyer but neither of these companies has any experience with processors. Samsung does memory products which is a long way from the logic intensive design of cpu's. These don't overlap at all in terms of manufacturing or design methodology. NVidia does chipsets and graphics and nVidia would be an obvious monopoly conflict due to AMD's ATI holdings. No doubt someone will mention IBM. However, IBM could never take over AMD's markets since IBM's use of AMD server processors would make it a competitor with its own customers. This would mean giving up the most profitable market segment of AMD's line and would probably impair volume with customers like HP, Gateway, and Dell that also produce servers. In other words, IBM would lose its current R&D support from AMD (the largest of all its partners), lose the technical support that AMD provides to Chartered, and lose AMD support for APM. The loss of volume would also make AMD processors cost more. Thus, IBM would turn a supporting partner and Intel hedge into a money losing niche processor division. Economically, this makes no sense at all. So, an AMD purchase is probably somewhat less likely than, say, striking oil in your backyard.

AMD managed to increase its revenues every quarter in 2007 as well as its Gross Margins. AMD finally managed to trim the quarterly loss to $164 Million. Now, I have to be honest and mention that AMD's current value per share of stock is worse than it was at the end of 2003. Adjusting for the change in number of shares AMD was comfortably higher by about $1.2 Billion but the $1.6 Billion goodwill charge has now brought this about $400 Million under. This is cause for concern if AMD's loss increases again in Q1. Supposedly AMD will break even in Q2. AMD's outlook though is oddly different from Intel's. Whereas Intel predicts a stagnant Gross Margin at a lower value than even 2004, AMD predicts that the current level of 44% will hold during the first half of 2008 but increase to 50% in the second half of 2008. So, why would this be? First of all, by Q3 AMD should finally have at least 2.8Ghz chips. This would cover nearly the entire range since it looks like Intel won't go higher than 3.16Ghz. There is also the question of 45nm from AMD. Statements made by AMD during the Earnings Call include:

Ruiz
“look forward to being able to ramp 45 nanometer aggressively in the second half of this year.”

Meyer
“Followed on in the second half of the year by 45 nanometer.”

Meyer
“we’ve got internal samples of our 45 nanometer microprocessors, we’re putting them through their paces currently and we’re on track to, the plans we talked about in the past which is to start our ramp in the first half of this year and ship revenue product in the second half of this year.”

I have seen some attempt to spin these statements to mean that AMD will begin producing 45nm in the second half with actual delivery in 2009. This, however, matches neither what AMD said nor common sense. Someone might mistakenly take Ruiz's comment to mean initial production however this is completely countered by Meyer's statement. Secondly, AMD first produced Brisbane samples in Q2 2006 and delivered product in Q4. If AMD has samples in Q1 then product delivery should by Q3. So, the most reasonable assumption is that AMD will begin production in Q2 and deliver some small amount in Q3. This amount is unlikely to have any effect on revenues but the Q4 volume should be more significant. Likewise, I doubt Intel's volume of Nehalem in Q4 will have any effect on revenues. So, adding in AMD's mobile and higher clocked quad cores in Q3 along with some 45nm in Q4 I could see an increase in Gross Margin in the second half. Finally, this begins to make sense. Right now, AMD overlaps the highest volume range and should increase its overlap during 2008. This would put pricing pressure on Intel with most likely some loss of ASP and this would offset the 45nm savings. On other hand, AMD's ASP's are very low so it can only move up. It also makes sense because 50% Gross Margin would still be significantly below Intel's 57%. This would be consistent with lower ASP's for AMD and higher costs due to a lower ratio of 45nm production. Some might also have noticed that AMD has scaled back its expected 2008 volume from 100 million units to 80-90 million units. This is undoubtedly due to the scale back in the FAB 38 ramp. This in turn has undoubtedly reduced AMD's target of 30% share for 2008.

I guess the bottom line is that as long as AMD can reach break even in Q2 it should be able to start gaining asset value again. And, it looks like AMD will finally get back to the processor revenues that it had in 2006. Judging from the current unit volume though I would say that AMD intends to start gaining again, especially in servers and Intends to get back to what it had in 2006. The really surprising thing is a comparison with 2001. AMD's volume share had been about 16% but this increased to 20% in 2001. With the release of the Northwood P4 and AMD's delay in getting to 130nm, AMD's share tumbled in early 2002. AMD left 2002 with roughly same 16% volume share that it had in 2000. It stayed at this level all through 2004 then creeped up to 18% in 2005. AMD's volume gain in 2006 was much more dramatic rising to 23%. If the 2002 pattern had been followed (as many predicted) then AMD should have lost this share again in 2007. When AMD's volume share tumbled in Q1 2007 many of these armchair experts gave themselves big pats on the back. However, they couldn't have been more wrong. AMD's volume share promptly rebounded and AMD has held onto its 2006 average for the last three quarters. What these self styled experts failed to take into consideration was that when the crash occured in 2002 AMD had only been at its high point for two quarters whereas when the crash came in 2007 it followed five quarters of good volume. In other words, the two are opposites. The high in 2001 was temporary whereas the dip in early 2007 was temporary. According to IDC AMD lost 0.4% unit share and is now at 23.1%. This is about equal to AMD's 2006 average. Overall, I expect AMD to gain back its 2006 server share and go above its previous mobile share while holding onto and increasing its current desktop share. This would mean an increase in unit share from the current 23% to perhaps 26% in 2008. This seems doable but falls quite a bit short of AMD's previous target of 30%.

Intel in 2008 should finally pull ahead of its previous processor revenue high of 2005. It's Gross Margins may not reach what they were but they will still be quite good and I'm sure Intel will continue increasing its cash and doing stock buybacks. By any account this should be good performance and a reasonable rate of growth. Pricing pressure in the second half of 2008 as AMD moves up above 2.6Ghz should produce some good values for buyers. We can look forward to seeing how well K10 scales and whether or not 45nm Shanghai produces any change in speed or any reduction in power draw. We should also be getting reviews of Nehalem in Q4.

Monday, November 05, 2007

Has Intel's Process Tech Put Them Leagues Ahead?

There has been a lot of talk lately suggesting that Intel is way ahead of AMD because of superior R&D procedures. Some of the ideas involved are rather intriguing so it's probably worth taking a closer look.

You have to take things with a grain of salt. For example, there are people who insist that it wouldn't matter if AMD went bankrupt because Intel would do its very best to provide fast and inexpensive chips even without competition. Yet these same people will, in the next breath, also insist that the reason why Intel didn't release 3.2Ghz chips in 2006 (or 2007) was because, "they didn't have to". I don't think I need to go any further into what an odd contradiction of logic that is. At any rate, the theory being thrown around these days is that Intel saw the error of its ways when it ran into trouble with Prescott. Then it scrambled and made sweeping changes to its R&D department, purportedly mandating Restricted Design Rules so that the Design staff stayed within the limitations of the Process staff. The theory is that this has allowed Intel to be more consistent with design and to leap ahead of AMD which presumably has not instituted RDR. The theory also is that AMD's Continuous Transistor Improvement has changed from a benefit to a drawback. The idea is that rather than continuous changes allowing AMD to advance, these changes only produce chaos as each change spins off unexpected tool interactions that take months to fix.

The best analogy of RDR that I can think of is Group Code Recording and Run Length Limited recording. Let's look at magnetic media like tape or the surface of a floppy disk. Typically a '1' bit is recorded as a change in magnetic polarity while a '0' is no change. The problem is that this medium can only handle a certain density. If we try to pack too many transitions too closely together they will blend and the polarity change may not be strong enough to detect. Now, let's say that a given magnetic medium is able to handle 1,000 flux transistions per inch. If we record this directly then we can do 1,000 bits per inch. However, Frequency Modulation puts an encoding bit between data bits to ensure that we don't get two 1 bits in a row. This means that we can actually put 2,000 encoded bits per inch and of this 1,000 bits is actual data. We can see that although FM expanded the bits by a factor of 2 there was no actual change in data density. However, by using more complex encoding we can actually increase density. By using (1,7) RLL we can record the same 2,000 encoded bits per inch but we get 1,333 data bits. And, by using (2,7) RLL we space out the 1 bits even further and can double the recording density to 4,000 encoded bits per inch. This increases our data bits by 50% to 1,500. GCR is similar as it maps a group of bits into a larger group which allows elimination of bad bit patterns. You can see a detailed description of MFM, GCR, and RLL at Wikipedia. The important point is that although these encoding schemes initially make the data bits larger they actually allow greater recording densities.

RDR would be similar to these encoding schemes in that while it would initially make the design larger it would eliminate problem areas which would ultimately allow the design to be made smaller. Also, RDR would theoretically greatly reduce delays. When we see that Intel's gate length and cache memory cell size are both smaller than AMD's and we see the smooth transition to C2D and now Penryn we would be inclined to give credit to RDR much as EDN editor, Ron Wilson did. You'll need to know that OPC is Optical Proximity Correction and that DFM is Design For Manufacturability. One example of OPC is that you can't actually have square corners on a die mask so this is corrected by rounding the corners to a minimum radius. DFM just means that Intel tries very hard not to design something that it can't make. Now, DFM is a good idea since there are many historical examples of designs from Da Vinci's attempts to cast a large bronze horse to the Soviet N1 lunar rocket that failed because manufacturing was not up to design requirements. There are also numerous examples from the first attempts to lay the Transatlantic Telegraph Cable (nine year delay) to the Sidney Opera House (eight year delay) that floundered at high cost until manufacturing caught up to design.

I've read what both armchair and true experts have to say about IC manufacturing today and to be honest I still haven't been able to reach a conclusion about the Intel/RDR Leagues Ahead theory. The problems of comparing the manufacturing at AMD and Intel are numerous. For example, we have no idea how much is being spent on each side. We could set upper limits but there is no way to tell exactly how much and this does make a difference. For example, if the IBM/AMD process consortium are spending twice as much as Intel on process R&D then I would say that Intel is doing great. However, if Intel is spending twice as much then I'm not so sure. We also know that Intel has more design engineers and more R&D money than AMD does for the CPU design itself. It seems that this could be the reason for smaller gate size just as much as RDR. It is possible that differences between SOI and bulk silicon are factors as well. On the other hand, the fact that AMD only has one location (and currently just one FAB) to worry about surely gives them at least some advantage in process conversion and ramping. I don't really have an opinion as to whether AMD's use of SOI is good idea or a big mistake. However, I do think that the recent creation of the SOI Consortium with 19 members means that neither IBM nor AMD is likely to stop using SOI any sooner than 16nm which is beyond any current roadmap. I suppose it is possible that they see benefits (from Fully Depleted SOI perhaps) that are not general knowledge yet.

There is at least some suggestion in Yawei Jin's doctoral dissertation that SOI could have continuing benefits. The paper is rather technical but the important points are that SOI begins having problems at smaller scale.

"we found that even after optimization, the saturation drive current planar fully depleted SOI still can’t meet 2016 ITRS requirement. It is only 2/3 of ITRS requirement. The total gate capacitance is also more than twice of ITRS requirement. The intrinsic delay is more than triple of ITRS roadmap requirement. It means that ultra-thin body planar single-gate MOSFET is not a promising candidate for sub-10nm technology."

The results for planar double gates are similar: "we don’t think ultra-thin body single-gate structure or double-gate structure a good choice for sub-10nm logic device."

However, it appears that "non-planar double gate and non-planar triple-gate . . . are very promising to be the candidates of digital devices at small gate length." But, "in spite of the advantages, when the physical gate length scales down to be 9nm, these structures still can’t meet the ITRS requirements."

So, even though AMD and IBM have been working on non-planar, double gate FinFET technology, this does not appear sufficient. Apparently this would have to be combined with novel materials such as GaN in order to meet the requirements. It then appears that it is possible for AMD and IBM to continue using SOI down to a scale smaller than 22nn. So, it isn't clear that Intel has any longterm advantage by avoiding SOI based design.

However, even if AMD is competitive in the long run that would not prevent AMD from being seriously behind today. Certainly when we see reports that AMD will not get above 2.6Ghz in Q4 that sounds like anything but competitive. When we combine these limitations with glowing reports from reviewers who proclaim that Intel could do 4.0Ghz by the end of 2008 this disparity seems insurmountable. The only problem is that the same source that says that 2.6Ghz Phenom will be out in December or January also says Fastest Intel for 2008 is 3.2GHz quad core.

"Intel struggles to keep its Thermal Design Power (TDP) to 130W and its 3.2GHz QX9770 will be just a bit off that magical number. The planned TDP for QX9770 quad core with 12MB cache and FSB 1600 is 136W, and this is already considered high. According to the current Intel roadmap it doesn’t look like Intel plans to release anything faster than 3.2GHz for the remainder of the year. This means that 3.2 GHZ, FSB 1600 Yorkfield might be the fastest Intel for almost three quarters."

But this is not definitive: "Intel is known for changing its roadmap on a monthly basis, and if AMD gives them something to worry about we are sure that Intel has enough space for a 3.4GHz part."

So, in the end we are still left guessing. AMD may or may not be able to keep up with SOI versus Intel's bulk silicon. Intel may or may not be stuck at 3.2Ghz even using 45nm. AMD may or may not be able to hit 2.6Ghz in Q4. However, one would imagine that even if AMD can hit 2.6Ghz in December that only 2.8Ghz would be likely in Q1 versus Intel's 3.2Ghz. Nor does this look any better in Q2 if AMD is only reaching 3.0Ghz while Intel manages to squeeze out 3.3 or perhaps even 3.4Ghz. If AMD truly is the victim of an unmanageable design process then they surely realized this by Q2 06. However, even assuming that AMD rushed to make changes I wouldn't expect any benefits any sooner than 45nm. The fact that AMD was able to push 90nm to 3.2Ghz is also inconclusive. The fact that AMD was able to get better speed out of 90nm than Intel was able to get out of 65nm could suggest more skill on AMD's part or it could suggest that AMD had to concentrate on 90nm because of greater difficulty with transistors at 65nm's smaller scale. AMD was delayed at 65nm because of FAB36 while Intel needs a fixed process for distributed FAB processing. Too often we end up with apples to oranges when we try to compare Intel with AMD. Also, we have to wonder why if Intel is doing so well compared to AMD with power draw then why did Supermicro just announce World's Densest Blade Server with Quad-Core AMD Opteron Processors.

To be honest I haven't even been able to determine yet if the K10 design is actually meeting the design parameters. There is a slim possibility that K10's could show up in the November Top 500 Supercomputer list. This would be definitive because HPC code is highly tuned for best performance and there are plenty of K8 results for comparison. Something substantially less than twice as fast per core would indicate a design problem. Time will tell.

Wednesday, September 19, 2007

The Top Developments Of 2007

It looks like both AMD and Intel have been as forthcoming as they are likely to be for awhile about their long range plans. The most significant items however have little to do with clock speeds or process size.

The two most significant developments have without doubt been SSE5 and motherboard buffered DIMM access. AMD has already announced its plan to handle motherboard buffered DIMMs with G3MX. This is significant because it means the end of registered DIMMS for AMD. With G3MX, AMD can use the fastest available desktop DIMMs with its server products. This is great for AMD and server vendors because desktop DIMMs tend to be both faster and cheaper than register DIMMs. This is also good news for DIMM makers because it would relieve them making registered DIMMs for a small market segment and allow them to concentrate on the desktop products. Intel may have the same thing in mind for Nehalem. There have been hints by Intel but nothing firm. I suppose Intel has reason to keep this secret since this would also mean the end of FBIMM in Intel's longterm plans. If Intel is too open about this it could make customers think twice about buying Intel's current server products which all use FBDIMM. So, whether this is the case with Nehalem or perhaps not until later it is clear that both FBDIMM and registered DIMMs are on their way out. This will be a fundamental boost to servers since their average DIMM speed will increase. However, this could also be a boost to desktops since adding the server volume to desktop DIMMs should make them cheaper to develop. This also avoids splitting the engineering resources at memory manufacturers so we could see better desktop memory as well.

SSE5 is also remarkable. Some have been comparing this with SSE4 but this is a mistake. SSE4 is just another SSE upgrade like SSE2 and SSE3. However, SSE5 is an actual extension to the x86 ISA. If AMD had been thinking clearer they might have called it AMD64-2. A good indication of how serious AMD is about SSE5 is that they will drop 3DNow support in Bulldozer. This clears away some bit codes that can be used for other things (like perhaps matching SSE4). Intel has already stated that they would not support it. On the other hand, Intel's statement means very little. We know that Intel executives openly lied about their intentions to support AMD64 right up until they did. And, Intel has every reason to lie about SSE5. The 3-way instructions can easily steal Itanium's thunder and Intel is still hoping (and praying) that Itanium will not get gobbled up by x86. Intel is also stuck in terms of competitiveness because it is too late to add SSE5 to Nehalem. This means that Intel would have to try to include it in the 32nm shrink which is difficult without making core changes. This could easily mean that Intel is behind in SSE5 until 2010. So, it wouldn't help Intel to announce support until it has to since supporting SSE5 now would only encourage development for an ISA extension that it will be behind in. Intel is taking the somewhat deceptive approach of working on a solution quietly while claiming not to be. Intel can hope that SSE5 won't become popular enough that it has to support it. However, if it does then Intel can always claim to be giving in to popular demand. It's dishonest but it is understandable for a company that has been painted into a corner.

AMD understands about being painted into a corner. Intel has had the advantage with MCM quad cores since separate dies mean both higher yields and higher clock speeds. For example, on a monolithic quad die you can only bin as high as the slowest core. However, Intel can pick and choose individual dies to put the highest binning ones together. Also, Intel can always pawn off a dual core die with a bad core as a lowly Conroe-L but it would be a much bigger loss for AMD to sell a quad die as a dual core. AMD's creative solution was the Triple Core announcement. This means that any quads with a bad core will be sold as X3's instead of X4's. This does make AMD's ASP look a bit better. I doubt Intel will follow suit on this but then it doesn't have to. For AMD, having an X4 knocked down to an X2 is a big loss but for Intel it just means having a Conroe knocked down to Conroe-L which is not so big. Simply put, AMD needs triple cores but Intel doesn't. On the other hand, just as AMD was forced to release a faster FX chip on the older 90nm process so too it seems Intel has been forced to deliver Tigerton not with the shiny new Penryn core but with the older Clovetown core. Tigerton is basically just Clovertown on a quad FSB chipset. This does suggest at least a bit of desperation since after working on this chipset for over a year Intel will be lucky if it breaks even on sales. To understand what a stumble Tigerton is you only have to consider the tortured upgrade path. In 2006 and most of 2007 Intel's 4-way platform meant Tulsa. Now we get Tigerton which uses the completely incompatible Caneland chipset. No upgrades from Tulsa. And, for anyone who buys a Tigerton system, oops, no upgrade to Nehalem either. In constrast, 4-way Opteron systems should be upgradable to 4-way Barcelona with just a BIOS update. And, if attractive, these should be upgradable to Shanghai as well. After Nehalem though, things become more even as AMD introduces Bulldozer on an incompatible platform. 2009 will without doubt be the year of new sockets.

For the first time in quite awhile we see Intel hitting its limits. Intel's 3.33Ghz demo had created the expectation of cool running 3.33Ghz desktop chips with 1600Mhz FSBs. It now appears that Intel will only release a single 45nm desktop chip in 2007 and it will only be clocked at 3.0Ghz. The chip only has a 1333Mhz FSB and draws a whopping 130 Watts. Thus we clearly see Intel's straining to deliver something faster much as AMD did recently with its 3.2Ghz FX. However, Intel is not straining because of AMD's 3.2Ghz FX chip (which clearly is no competition). Intel is straining because of AMD's server volume share. In the past year, AMD's sever volume has dropped from about 25% to only 13%. Now with Barcelona, AMD stands to start taking share back. There really isn't much Intel can do to prevent this now that Barcelona is finally out. But any sever chip share that is lost is a double blow because server chips are worth about three times as much as desktop chips. This means that any losses will hurt Intel's ASP and boost AMD's by much more than a similar change in desktop volume would. So, Intel is taking its best and brightest 45nm Penryn chips and allocating them all to the server market to try to hold the line against Barcelona. Of the 12% that Intel has gained it is almost certain to lose half back to AMD in the next quarter or two, but if it digs in, then it might hold onto the other half. This means that the desktop gets the short end of the stick in Q1 2008. However, by Q2 2008, Intel should be producing enough 45m chips to pay attention to the desktop again. I have to admit that this is worse than I was expecting since I assumed Intel could do a 3.33Ghz desktop chip by Q1. But now it looks like 3.33Ghz will have to wait until Q2.

AMD is still a bit of a wild card. It doesn't appear that they will have anything faster than 2.5Ghz in Q4 but 3.0Ghz might be doable by Q1. Certainly, AMD's demo would suggest a 3.0Ghz in Q1 but as we've just seen, demos are not always a good indicator. Intel's announcement that Nehalem has taped out is also a reminder that AMD has made no such announcement for Shanghai. AMD originally claimed mid 2008 for Shanghai and since chips normally appear about 12 months after tapeout we really should be seeing a tapeout announcement very soon if AMD is going to release by Q3 2008. There is little doubt that AMD needs 45nm as soon as possible to match Intel's costs as Penryn ramps up. A delay would seem odd since Shanghai seems to have fewer architecture changes than Penryn. AMD needs a tapeout announcement soon to avoid rumors of problems with its immersion scanning process.

Thursday, August 23, 2007

2008 And Beyond

2007 is far from over but it seems that lately people prefer to talk about 2008. Perhaps this is because AMD is unlikely to get above 2.5Ghz with K10 and Penryn will only have a low volume of about 3%. I suppose this is not a lot to get excited about. So, we are encouraged to cast our gaze forward but what we see is not what we might expect.

AMD's server chip volume has dropped considerably since last year. So, there is little doubt that this trend will reverse in Q3 and Q4 of 2007 with Barcelona. This is true because even at lower clock speeds, Barcelona packs considerably more punch than K8 Opteron at similar power draw. The 2.0Ghz Q3 chips should replace around half of AMD's current Opterons and faster 2.5Ghz chips replacing even the fastest 3.0Ghz K8 Opterons in Q4. This should leave Intel with two faster server chip speeds in Q4 with this most likely falling to a single speed in Q1 08. However, Intel may be able to pull farther ahead in Q2 08. I'm sure this will be confusing to those who are comparing the Penryn launch with Woodcrest last year and assuming that the highest speed grades will be released right away. The problem with this view is that Penryn is leading 45nm in Q4 of this year whereas Woodcrest did not lead 65nm in 2006. Instead, Woodcrest was six months behind Presler which went into 65nm production in October 2005 and launched in December 2005. This explains why Woodcrest was able to hit the ground running and launch at 3.0Ghz. June 2006 was six months after 65nm Presler in December 2005. Taking this as the pattern for 45nm would mean top initial speeds wouldn't be available until Q2 2008. This seems true since Intel has been pretty quiet about Q1 08 release speeds. If the market expands in early 2008, Intel should get a boost as AMD feels the pinch in volume capacity caused by the scale down at FAB 30 and the increased die size of quad core K10. This combines with Intel's cost savings due to ramping 45nm to put Intel at its greatest advantage. However, by the end of 2008, this advantage will be gone and Intel won't see any new advantage until 2010 at the earliest.

To understand why Intel's window of advantage is so small you need to be aware of the differences in process introduction timelines, ramping speeds, base architecture speed, and changing die size advantages. A naiive assumption would be that: 1.) Intel's timeline maintains a process launch advantage over AMD, 2.) that Intel transitions processes faster, 3.) that Penryn is considerably faster than Conroe and that Nehalem is considerably faster than Penryn, and 4.) that Nehalem maintains Penryns's die size advantage. However, each of these assumptions would be incorrect.

1.) Timeline

Q2 06 - Woodcrest
Q3 07 – Barcelona Trailing by 5 quarters.

Q4 07 - Penryn
Q3 08 – Shanghai Trailing by 3 quarters.

Q4 08 - Nehalem
Q2 09 – Bulldozer Trailing by 2 quarters.

Q4 09 - Westmere
Q1 10 - 32nm Bulldozer Trailing by 1 quarter.

Intel's Tick Tock timeline is excellent but AMD's timeline steadily shortens Intel's lead over the next two and a half years. This essentially means that the dominance that C2D enjoyed for more than a year will not be repeated. I suppose it is possible that 45nm will be late but AMD continues to say that it is on track. The main reason I am inclined to believe them is the die size. When AMD moved to 90nm they only had a small shrink in die size at first and then they later had a second shrink. AMD only reduced Brisbane's die size to 70% and nine months later AMD could presumably do a second shrink. But they aren't; Barcelona shows the same 70% reduction as Brisbane. This suggests to me that AMD has skipped a second die shrink and is concentrating on the 45nm launch. I'm pretty certain that if 45nm were going to be late that we would be seeing another shrink of 65nm as a stopgap.


2.) Process Transition

Most people who talk about Intel's process development only know that Intel launches a process sooner than AMD. However, the amount of time it takes Intel to actually field a new process is also important. Let's look at Intel's 65nm history starting with an Intel Presentation concerning process technology. Page 2:

Announced shipping 65nm for revenue in October 2005

CPU shipment cross-over from 90nm to 65nm projected for Q3/06


And, from Intel's website, 65-Nanometer Technology:

Intel has been delivering 65nm processors in volume for over one year and in June 2006 reached the 90-65nm manufacturing "cross-over," meaning that Intel produced more than half of total mobile, desktop and server microprocessors using industry-leading 65nm process technology.

So, we can see that Intel did quite well and even beat its own projection by reaching crossover in late Q2 instead of Q3. October 2005 to June 2006 would be eight months to 50% conversion. For AMD, the INQ had a rumor for shipping in October and we know it officially launched December 5th 2006. Let's assume that this is true since it matches with Intel's October revenue shipping date with a December release in 2005. The AMD Q1 2007 Earnings Transcript from April 19th 2006 says:

100% of our fab 36 wafer starts are on 65 nanometer technology today

October 2006 to April 2007 would be 6 months. So, this would mean that AMD made a 100% transition in two months less than it took Intel to reach 50%. Intel's projection of 45nm is very similar with crossover not occuring until Q3 08. What this means is that even though Intel launches 45nm with a headstart in Q4 07, AMD should be completely caught up by Q1 09.


3.) Base Architecture Speed

Intel made grand claims of a 25% increase in gaming performance (40% faster for 3.33Ghz Penryn versus 2.93Ghz Kentsfield). However, according to Anandtech's Wolfdale vs. Conroe Performance review, Penryn is 4.81% faster while HKEPC gets 5.53% faster. A 5% speed increase is similar to what AMD got when it moved from 130nm to 90nm. The problem that I see is not with Intel's exageration but that Nehalem seems to use the same core. In fact, other than HyperThreading there seems to be no major changes to the core between Penryn and Nehalem. The main improvements with Nehalem seem to be external to the core like an Integrated Memory Controller, point to point communications, L3 cache, and enhanced power management. The real speed increases seem to come primarily from GPU processing and ATA instructions however like Hyperthreading these are not going to make for significant increases in general processing speed. And, since Westmere is the same core on 32nm this means no large general speed increases (aside from clock increases) for Intel processors until 2010 at the earliest. I suppose this then leaves the question of whether AMD will get a larger general speed increase with Bulldozer. Presumably if AMD can manage it they could then pull ahead of Nehalem. Both Intel and AMD are going to use GPU's on the die and both are going to go to more cores. Nehalem might get ahead of Shanghai since while both can do 8 cores Nehalem can also do HyperThreading. But Bulldozer moves back ahead again by allowing 16 actual cores. At the moment it is difficult to imagine a desktop application that could effectively use 8 cores, much less 16 but who knows how it will be in two years.


4.) Die Size

For AMD the goal is to get through the first half of 2008 because the game looks quite different toward the end of 2008. By the time Nehalem is released Intel will already have gotten most of the benefit of 45nm while AMD will only be starting. Intel will lose its small die size MCM advantage because Nehalem is a monolithic quad die like Barcelona. Intel only got a modest shrink of 25% on 45nm and so far has only gotten a 10% reduction in power draw so AMD can certainly stay in the game. It is also a certainty that Nehalem will have a larger die size than quad Penryn. This will be true because Nehalem will have to have both an Integrated Memory Controller and the point to point CSI interface. Nehalem will also add L3 cache. It would not be surprising if the Nehalem die is larger than AMD's Shanghai die. The one positive for Intel is that although yields will be worse with a monolithic die, their 45nm process should be mature by then. However, AMD has shown considerably faster process maturity so yields should be good on Shanghai in Q1 09 as well.

An Aside: AMD's True Importance

Finally, I have to say that AMD is far more important than many give them credit for. I recall a half-baked editorial by Ed Stroligo A World Without AMD where he claimed that nothing much would change if AMD were gone. This notion shows a staggering ignorance of Intel's history. The driving force behind Intel's advance from 8086 to Pentium was Motorola whose 68000 line was initially ahead. It had been Intel's intention all along to replace x86 and Intel first tried this back in 1981 with iAXP 432. It's segmented 16MB addressing looked pretty good compared to 8086's 1MB segmented addressing. However, it looked a lot worse than 68000's flat 16MB addressing which had been released the year before. The very next year iAXP 432 became the Gemini Project which then became the BiiN company. IAXP 432 continued in development with the goal of replacing x86 until 1989. However, this project could not keep up with the rapid pace of x86 as it struggled to keep up with each generation of 68000. When Biin finally folded, a stripped down version of iAXP 432 was released as the embedded i960 RISC processor. Interestingly, as the RISC effort ran into trouble Intel began working on VLIW and when BiiN folded in 1989 Intel released its first VLIW procesor, i860. HP began work on EPIC the same year and five years later, Intel was commited to EPIC VLIW as an x86 replacement.

In 1995 Intel introduced Pentium Pro to take on the established RISC processors and grab more share of the server market. The important point though is that there is no indication that Intel ever intended Pentium Pro to be used on the desktop. We can infer this for a couple of reasons. First, Itanium had been in development for a year when Pentium Pro was introduced and an Itanium release was expected in 1998. Second, with Motorola out of the way (68000 development ended with 68060 in 1994), Intel was not expecting any real competion on the desktop. AMD and Cyrix were still making copies of 80486 so Intel had only planned some modest upgrades to Pentium until Itanium was released. However, AMD released K5 which thoroughly stunned Intel. Although K5 was not that fast it did have a RISC core (courtesy of AMD's 29050 RISC processor) which put K5 in the same class as Pentium Pro and a generation ahead of Pentium. Somehow AMD had managed the impossible and had skipped the Pentium generation. So, Intel went to an emergency plan and two years later released a cost reduced version of Pentium Pro for the desktop, Pentium II. The two year timeline indicates that Intel was not working on a desktop version previous to K5's release. Clearly, we owe Pentium II to K5.

However, AMD purchased Nexgen and released the powerful K6 (which also had a RISC core) just two years later meaning that it arrived at the same time as PII. Once again Intel was forced to scramble and release PIII two years later. We owe PIII to K6. But, AMD had been hard at work on a K5 successor and with the added technology from K6 and some Alpha tech it released K7. Intel was even more shocked this time because K7 was a generation ahead of Pentium Pro. Intel was out of options so it was forced to release the experimental Williamette processor and then follow up with the improved Northwood two years later. We owe P4 to K7. That P4 was experiemental and never expected to be released is quite clear from the pipeline length. The Pentium Pro design had a 14 stage pipeline which was reduced to 10 stages in PII and PIII. Interestingly Itanium also used a 10 stage pipeline. However, P4's pipeline was even bigger than the original Pentium Pro's at 20 stages. Itanium II has an even shorter pipeline at 8 stages so it is clear that Intel does not prefer long pipelines. We can then see that P4 was an aberration caused by necessity and Prescott at 31 stages was a similar design of desperation. Without K8 there would be no Core 2 Duo today and without K10 there would be no Nehalem.

There is no doubt whatsoever that just as 8086's rapid advance against competition from Motorola 68000 stopped the iAXP 432 and shutdown Biin, Intel's necessity of advancing Pentium Pro rapidly on the desktop stopped Itanium. Intel already had experience with VLIW from i860 and would have delivered Merced on schedule in 1998. Given Itanium's speed it could have been viable at as little as 150Mhz. However, Pentium II was already at 450Mhz in 1998 with faster K7 and PIII speeds due the next year. The pace continued rapidly going from Pentium Pro's 150Mhz to PIII's 1.4Ghz. Itanium development simply could not keep up and the grand plans of 1997 for Itanium to become the dominant processor fell apart. The pace has been no less relentless since PIII and Itanium has been kept in a niche server market.

AMD is the sole reason why today Itanium is not the primary processor architecture. To suggest that nothing would change if AMD were gone is an extraordinary amount of self delusion. Intel would happily stop developing x86 and would put its efforts back into Itanium instead. The x86 line is also without any serious desktop replacement. Alpha, MIPS, and ARM stopped being contenders long ago. Power was the last real competitor but it fell out of the running when its desktop chips couldn't keep up and were dropped by Apple. This means that without AMD, Intel's sole competition for desktop processors is VIA. And, just how far behind is VIA? No AMD would mean higher prices and slower development and the eventual phase out of x86. Of course, I guess people can always hope that Intel has given up its goal of more than a quarter century of dropping the x86 line and moving the desktop to a completey proprietary platform.

Friday, August 17, 2007

2007: The Second Half

Amid all the rumblings and rumors there signs of fundamental differences between this year and last. In almost every aspect of processors AMD and Intel have swapped places. This has left a virtual vacuum of analogy for AMD and Intel supporters alike since both are reluctant to compare their favorite to the competition. The situation today is not exactly the same but some comparisons do provide a view of where things are likely to go.

We can add up the various ways that Intel and AMD have swapped places and there are quite a few. In early 2006, AMD's K8 was the undisputed leader ahead of Intel's Presler and Yonah offerings. Today, C2D is the undisputed leader ahead of AMD's Opteron and Athlon 64 offerings. In late 2006, Intel introduced quad core which AMD has taken nearly a year to match. Today, AMD is ready to offer native quad core which it will take Intel about a year to match. In early 2006, Intel previewed the native dual core 2.93Ghz Conroe which looked great and then it was a matter of waiting for Intel to actually get them out the door in volume. Today, AMD has previewed native quad core 3.0Ghz K10 which looks great and once again it is a matter of waiting for AMD to get them out the door in volume. In 2006, Intel was recovering from revenue shocks caused by AMD's K8. Today AMD is recovering from revenue shocks caused by Intel's C2D. In 2006, Intel introduced a new architecture that was far ahead of its previous generation offerings while AMD was only able to offer secondary upgrades such as small clock increases, virtualization, and faster memory speeds. Today, AMD is offering a new architecture that is far ahead of its previous generation offerings while Intel is only able to offer secondary upgrades such as small clock increases, SSE4, and faster bus speeds.

It really is remarkable how similar each company's situation is to its competitor's last year. This is most fundamentally true on the desktop. I suppose Intel supporters would point out that AMD is not likely to take the top performance spot in Q4 when Phenom is launched as Intel did when Conroe was launched. That is true. However, I suppose AMD supporters could point out that AMD was never in the heat and power crunch that Prescott was. Mobile is the most fundamentally different. Intel took mobile by storm when it launched the Centrino platform and since that time AMD has only been slowly chipping away with Turion. Merom was nearly the opposite of Conroe. While Conroe added tremendous value to the desktop as it replaced the sagging P4 line, Merom on the other hand actually had worse power draw than Yonah. Intel finds that having conquered the battery life and wireless LAN issue years ago that it has no place left to take mobile to gain an advantage. Turion has only had a small effect on Intel's mobile share but AMD should be fully competitive in 2008 with Griffin and Puma. It also looks like most of Intel's tweaks with Penryn are to try to stave off the coming attack from K10 Opteron. Intel is putting up a good fight with lower power draw, more cache, and faster FSB but it won't be enough. The fact is that when you've taken back as much server share as Intel has the only place left to go is down. SSE4 could be a big boost in HPC however Intel has already made most of its gains in the HPC low range with Woodcrest so Penryn would most likely be an upgrade to existing systems. SSE4 could be a boost in the top range but currently Intel has little presence there.

With Intel certain to have small losses in server share and no real change from the previous situation in mobile that leaves the desktop as the main battleground. AMD's average volume share in 2005 was 18% up noticeably from about 16.5% average for the previous several years. AMD's average volume share in 2006 was 23% and even though Intel has been fighting hard it remains at 23% in Q2 07. The actual price cuts have been a lot less than most people imagine. Intel's overall ASP is only down 16% from the nearly steady value of about $99 that it had been for three quarters. AMD's drop is similarly down 17% from the previous three quarter average of $60. AMD has had a desktop ASP drop of 42% since Q1 06 while Intel has dropped 38% in the same time. In Q2 07 AMD's desktop ASP was $49 versus Intel's $83. AMD's desktop ASP is substantially lower than Intel's but if it remains steady AMD will make more money as its margins improve with cost savings from 65nm. Although Intel's reorganization has so far only brought tiny changes in cost reduction it could see more in the 2H of 2007.

This does bring up the question of whether Intel will be able to bring additional price pressure to bear against AMD. Intel's Q2 07 earnings suggest that Intel reached its lower limit in pricing in Q2 and that it would need additional cost savings to be able to price lower. This plus the Q2 07 reductions in ASP for both server and mobile make it unlikely that we will see much in the way of lower prices during the rest of 2007. However, Intel is almost certain to resurrect this tactic in some fashion in 2008 as its ramping 45nm production reduces costs again. This should be interesting since AMD's ramping K10 desktop production should raise its desktop ASPs. It wouldn't surprise me to see a substantial bump of AMD's desktop ASP to $60 with a small cut of Intel's down to $79. This is possible if everything goes well and Intel still wants to keep prices down. Otherwise I would expect Intel to pull its desktop ASP back up to its preferred level of $99 and AMD to increase its to a preferred level of $70. This will likely be dependent on Intel's flash spin-off not being a $300 Million a quarter drain and AMD's new chipset division earning a profit. Mostly this means that Q4 07 will be more of a skirmish than the major battle that was expected. Presumably this will become a genuine battle in 2008 as Intel ramps Penryn while AMD ramps K10.

Some people seem to assume that Intel will ramp 45nm quickly and have large volumes available in 2007. However, the following ramp graph from Intel shows that 45nm will only be about 3% of production before the end of 2007. 45nm won't be a significant desktop volume for Intel until Q2 08 with crossover occuring the following quarter. Again, this why 2008 will be the real battle.