Oracle Corporation (ORCL)
NYSE: ORCL · Real-Time Price · USD
148.56
+0.95 (0.64%)
At close: Sep 21, 2026, 4:00 PM EDT
149.08
+0.52 (0.35%)
After-hours: Sep 21, 2026, 7:59 PM EDT
← View all transcripts

Status Update

Oct 8, 2013

Operator

Good day, ladies and gentlemen. Welcome to the Oracle Big Data Primer webinar call. At this time, all participants are in a listen-only mode. Later, we will conduct a question and answer session, and instructions will follow at that time. If anyone should require assistance during the conference, please press star then zero on your touchtone telephone. As a reminder, this conference call is being recorded. I would now like to introduce your host for today's conference, Ms. Shauna O'Boyle. Ms. Shauna, you may begin your conference.

Shauna O'Boyle
Senior Manager of Investor Relations, Oracle

Thanks. Hello, everyone. Thank you for joining us today as part of our ongoing educational speakers series hosted by Oracle. I am Shauna O'Boyle, Senior Manager of Investor Relations, and today is Tuesday, October eighth, two thousand thirteen. Joining us today is Oracle Executive Senior Vice President, Andrew Mendelsohn, and Equity Research Analyst Brendan Barnicle of Pacific Crest. Today, Andy will be discussing big data. He will not be discussing any data that is not already publicly available. At the conclusion of Andy's presentation, we will turn the webcast over to Brendan, who will moderate the question and answer portion of the call. You may submit questions at any time during the presentation by typing your question in the Q&A box in the lower part of your screen. Please keep in mind that we will not comment on business in the current quarter.

As a reminder, the matters we will be discussing today may include forward-looking statements, as such, are subject to the risk and uncertainties that we will discuss in detail in our documents filed with the SEC, specifically the most recent reports on Form 10-K and 10-Q, which identify important risk factors that may cause actual results to differ from those contained in forward-looking statements. You're cautioned not to place undue reliance on these forward-looking statements, which reflect our opinions only as of the date of this presentation. Please keep in mind that we're not obligating ourselves to revise, update, or publicly release the results of any revisions of these forward-looking statements in light of new information or current events. An authorized recording of this conference call is not permitted. I would now like to introduce Andrew Mendelsohn.

Andrew Mendelsohn
Executive Senior Vice President, Oracle

Thanks, Shauna. Good morning, everybody. I'm going to give a very short, about 15-minute chat about big data. Then we'll take questions. I thought I'd first start talking about what is big data. There have been all kinds of people talking about what big data is. There's a famous three Vs or four Vs of big data, and all the various new startup companies that have anything to do with information are claiming they're big data. What is the real kernel of what's going on here? The real kernel is that what we're talking about here is it's all about analytics. People have been doing analytics for 30, 40 years with information systems. What we're talking about here is moving to the next generation of those analytics, which we're calling big data analytics.

There are two key transformations going on in the industry that are driving the big data trend. Number one is really about different kinds of data. Traditionally, analytic systems have been looking at data from companies' operational systems. For example, if you're a big retailer, you're looking at data from your retail sales, the sales of information from products at your retail stores. If you're a telco, you're looking at your call data records that record all the phone calls and how long they've been there and what the charges are and things of that sort. That's analytics on your operational data. What people are talking about with big data is looking at more kinds of data than traditionally they've been looking at. Some of these data sources are from inside the company, things like looking at documents, more unstructured data, voice, video.

As we move to the Internet of Things, people are looking at sensor data sources as well. On the Internet, there's also a huge amount of data, of course, on the Internet, and companies are looking at, well, is there some way I can extract useful information out of social media? Can I find information about my customers again, and what they're saying on social media about my company? Are they saying something good? Are they saying something bad? There's a lot of bloggers out there saying all kinds of things that are potentially of interest to companies. People are very excited about the possibility of looking at this data, getting more information, especially about their customers, and using that to raise the potential revenue of their companies by better marketing their customers, better upselling them, et cetera.

I think that's the first big thing, moving from operational data sources to these broader internal and Internet sources. The next big thing going on is new kinds of analytics are being done. Traditionally, people did what they call OLAP analytics, slice and dice of data, mostly historical data from the operational systems. We've been moving over the last few years to more predictive analytics, data mining techniques, for example, to look at what's going to happen in the future. As we move to these broader kinds of data sources, people are looking at doing analytics against those kind of data sources, text, spatial. Graph analytics has become very popular as you look at social networks. People want to know or inquiries about who is a friend of who, and based on that, can they do some kind of interesting marketing or targeted advertising, et cetera.

We've also been moving to new kinds of analytic tools. R has become very popular as a development tool for doing analytics processing. Of course, we've been in memory, in a lot of phases to do in-memory analytics, in-memory databases, et cetera. I think those are the two big key drivers of what's going on in big data. What are we doing at Oracle to deal with this new world? Well, Oracle, of course, has for many years been the market leader in analytics. We have the world's leading data warehousing technology, BI technology with our database and our BI products as well. What we're doing as we move forward to the big data space is we, of course, are trying to build up a platform and a set of solutions that deal with big data.

We need to be able to acquire, organize, discover, and analyze big data. As we move forward, in addition to the Oracle massively parallel relational database for doing analytics against big data, we're also now moving to utilize Hadoop as a platform in our big data environments. We've also come out with Oracle's NoSQL Database that has a value in the ingestion phase of big data. Of course, we're continuing to evolve our big data tools as well. One of the big things we're doing at Oracle that I think is very important in the big data space is we are in the business of delivering what we call engineered systems.

These are combinations of hardware and software that are integrated together to deliver a great off-the-shelf experience for customers where they can order our products or our hardware and software and get great time to value, where they can very quickly build systems. We've done that in the big data space. We have our Big Data Appliance for running the Hadoop processing on our Oracle NoSQL Database as well. We have Oracle Exadata, which is a great platform for doing massively parallel analytics. Then, of course, we've finally added Exalytics, which is our platform for our BI tools and our Endeca data discovery tool.

Finally, on the top, Oracle is, of course, in the applications business, we are also in the business of delivering horizontal and vertical applications and solutions to our customers, and we will continue doing that on top of this stack of big data technologies as well. That's our key strategy. Now let's move on sort of a little closer to the products that we actually are doing. In this picture, we're showing, of course, on the left side, the data sources that we talked about, all kinds of data sources, both operational and new kinds of data sources like documents and blogs and information off the Internet, et cetera.

What we're showing that's a little different now is, in the past, what you would do is you'd take this information, you might put it into various files and do staging operations, do ETL operations that transform the data before moving it into a data warehouse. Now, what we're showing in this next-generation platform is using Hadoop and Hadoop's HDFS distributed file system as a way of ingesting large amounts of this data. Then we will use the Hadoop MapReduce platform to do some batch processing and ETL transformations against that data, sift through the data, look for interesting tidbits before we then use our big data connectors and load that into the Oracle Exadata data warehouse. We also show Oracle's NoSQL Database here as well.

NoSQL databases are also very good at ingesting information rapidly and also, again, doing some operations against it, then that data can be also, again, moved into a data warehouse if necessary. On top of this basic platform of Hadoop and the Oracle Database, we have our analytics engines. We have in the database our advanced analytics capability, which includes predictive analytics and R processing. We are also making this kind of analytics, especially R, available on the Hadoop platform as well. Finally, on top, we show Oracle's business analytics tools like our BI EE tool , our Endeca data discovery tool. Those tools also can be run against data both in the Hadoop HDFS file system and in the Oracle Database to deliver visualizations and analytics against the data. Okay, let's go to the next slide.

As I mentioned, a big part of our strategy is our engineered system. In this next slide, I just sort of show how those engineered systems fit into the big data solution we just talked about. Of course, the Oracle Big Data Appliance is our engineered system for running both Hadoop and for running the Oracle NoSQL Database. We are the first vendor, by the way, that's producing an engineered system optimized for running Oracle NoSQL or any other NoSQL database for that matter. Oracle Exadata, of course, is our platform of choice for running big, massively parallel data warehousing, for doing your interactive analytics. Oracle Exalytics is our engineered system for running our BI analytics foundation, and that includes BI EE and the Oracle Endeca Information Discovery product, and also our Oracle Essbase engine as well. Okay, let's move on.

What I want to do at this point is drill down a little bit on sort of or talk about how Hadoop and Oracle databases relate to each other. I think there's a lot of confusion out there about what Hadoop is good for and what massively parallel relational databases are good for. The key thing to understand, and what we are doing in this platform, is that in order to have a big data solution, you need both. If you talk to even the people who are the biggest early advocates of Hadoop, who are now trying to do analytics, what they've decided is Hadoop is a great platform for ingesting large amounts of data at very low cost per terabyte. It's a great platform for doing some analytics on that data, but it's more batch processing analytics. What does that mean?

Well, it means if you have a data scientist who's sitting in front of your terminal, and he's asking questions, trying to understand the business and trying to come up with great ideas for raising revenue, better marketing, advertising, et cetera, he wants interactive response. He wants to send in a query and get a response back in a few seconds. That's not really what Hadoop was designed for. Hadoop is designed to crank out scalable batch processing execution, and you'll get maybe tens of minutes or an hour response to those kind of queries. What they want is snappy response, and that's what you need a massively parallel relational database to do, and that's what these guys are doing. They'll use Hadoop for ingestion and for doing some big batch processing analytics against big data.

They'll move a subset of the data that they want to do further analytics on into their massively parallel relational database, in this case, Exadata. That's where they'll do their interactive analytics against the data using, of course, rich SQL language that we provide with Oracle. The other thing to note is that although there are some SQL tools available in the Hadoop environment, they're very primitive and raw, Hive, Pig, et cetera. Then, of course, you can code in Java as well. On the massively parallel relational database side of the world, what people do is they code in SQL, for the most part. SQL is a very expressive and productive language. A couple lines of SQL is equal to hundreds of lines of code in Java.

It's also much more efficient at processing as well, it's much faster and requires much fewer computing resources to get a given job done. People also like the fact that the relational databases are just much more productive environments than Hadoop is today. We also mentioned R. R is, of course, something that we made available as a statistical programming language and predictive analytics language in both the Exadata platform, and we're now also making available on the Hadoop platform as well. Let's go to the next slide. Here, this is a slide taken out of my keynote from Oracle OpenWorld. You're all welcome to go to oracle.com and take a look at the keynote there I did in Oracle Database 12c.

What this did was, it's sort of the end result of an example we went through where we showed, let's say you want to do a very common example that people talk about in big data, which is looking for fraud in a banking system of some sort. We wrote the application two ways. We wrote it using Java on Hadoop using MapReduce, we also wrote it using SQL extension we call SQL Pattern Matching, which is a new part of the SQL language that we've implemented in the 12c version of our database. We just measured two key metrics here. One is, how many lines of code does it take to solve this problem? What you see here, it took over 650 lines of code using Java MapReduce versus, I think it's on the order of 15 lines of code using SQL.

Number one, SQL is much, much more productive than having to code at a much more primitive level using, in this case, Java MapReduce. We also wanted to show we can run SQL in a very high-performance, massively parallel fashion. In this example, the runtime also of the SQL version of the analytics was much, much less, less than 10 seconds. While the runtime on Hadoop was over 70 seconds. We actually ran the SQL on, I think, a couple processors, and Hadoop was on an 18-node cluster. The key message here is relational databases are constantly moving the bar and getting faster and faster.

I think people who are thinking that, oh, we're just going to put a little SQL engine on Hadoop and catch up in a couple of years to what relational databases are doing and have built over the last 20 years for doing high-performance, massively parallel SQL querying, I think are being a little optimistic about how soon they're going to get to parity there. Okay. Let's go to the next slide. Here, I just want to mention one thing that we're doing in Database 12c around in-memory processing. For analytics, the relational database engines have been producing very high-performance, massively parallel relational database engines for many years that are very good at cranking through, crunching through terabytes and petabytes of information.

There's been a sort of a breakthrough in the last few years in looking at column store technologies and, in particular, now in-memory column store technologies for making analytics even faster. In Database 12c, we just announced at Oracle OpenWorld a few weeks ago that we are adding this in-memory column store technology to Database 12c. This, again, is going to give us another big leap forward in analytic processing in the relational database, in this case, Oracle Relational Database. This, of course, is going to be very exciting for customers doing analytics against big data and data warehousing. Again, we're sort of raising the bar of what relational databases can do here. We're not standing still. Relational databases are moving forward very aggressively into the analytics space further.

Again, people who are building SQL engines from scratch, again, have to think about adding even more technologies than they're thinking of to sort of match the capabilities of these relational databases. Okay. Let's go to the next slide. What are some of the key differentiators for what Oracle's doing in the big data space versus other competitors? I think the big thing we are doing here is we are giving customers an integrated platform and an engineered platform. A customer can just come to Oracle and say, "Okay, I want a big data platform." They would order Oracle Big Data Appliance, Oracle Exadata platform. Those two platforms or engineered systems can be very easily integrated together. We have both hardware integration, be it by using InfiniBand networking technology across both platforms for making it very efficient to move information back and forth.

We have software integration that we call our connectors that tie together the Hadoop platform with the Oracle Database platform. For example, one of the connectors, our SQL connector, lets Oracle SQL reach out into Hadoop HDFS and run SQL queries against the HDFS data. Another example connector is our loader that lets you very efficiently move data from HDFS into the Oracle Exadata database. We have, of course, our whole array of technologies in Exadata that make it a great platform for doing big data analytics. Next big differentiation that I just mentioned, we are adding very high performance in-memory columnar processing into our already very powerful relational database engine in Oracle. That's going to make us even much more outstanding interactive platform for doing analytics against big data.

Another big part of what we're doing here is Oracle has a huge ecosystem around it of developers and ISVs and SIs who know and love the Oracle platform. They have huge sets of skilled consultants who know how to manage the platform. They know how to develop against it. We are sort of building on top of that as we go into the big data space. Finally, we are giving a complete solution to our customers. If you buy our Oracle Big Data Appliance, you buy Oracle Exadata and the connectors between those two. If you have any problem, of course, you just call up Oracle. We support you top to bottom from hardware to software with any issue you have. Of course, we can provide you all the consulting help you need as well on top of that. Okay. Let's just close.

Of course, we have huge numbers of customers using our big data platform. I'll just mention a couple here. UPMC is University of Pittsburgh Medical Center, which is one of the leading medical research centers in the fields of genetics and other health sciences. They are a huge Oracle customer for big data and big user of Oracle Exadata. For example, let's go through here. SoftBank, big telco, huge data warehousing user. SoftBank is a big telco in Japan. They've actually also recently just been moving into the U.S. They actually moved all their warehousing technology off Teradata onto Oracle Exadata several years ago. They're a very successful customer. Thomson Reuters is one of our big data customers. They are using the Oracle Big Data Appliance and Oracle Exadata in their big data processing. StubHub is part of eBay. They do ticket reselling.

They're using Oracle's R technology that I mentioned earlier for doing statistical analysis and predictive analytics. With that, I think we'll move on to the Q&A section.

Shauna O'Boyle
Senior Manager of Investor Relations, Oracle

Thank you, Andy. Before I turn the call over to Brendan for the question and answer portion of the call, please let me remind our listeners that you can submit questions at any time during the presentation by typing your question in the Q&A box at the lower part of your screen. Brendan?

Brendan Barnicle
Equity Research Analyst, Pacific Crest

Thanks so much, Shauna, and thanks, Andy. Andy, I wanted to follow up on some of those customer references you just gave. I was wondering where you're seeing your customers most frequently use your big data solutions right now, if there is a most frequent use case.

Andrew Mendelsohn
Executive Senior Vice President, Oracle

I think the most popular use cases these days are in financial services and telco. The big banks are definitely very interested in a lot of what we're talking about. They are very interested, of course, in looking at data from social networks to see if they can get information about their customers. They're also interested in using that data to see if they can use it to help in fraud detection. Telcos are another big vertical that we see a lot of interest. A big problem in telco, of course, are customers leaving one telco to go to another. They call this churn. Churn analytics is a big part of what they're doing. They're actually looking at using graph analytics for doing that kind of processing.

Brendan Barnicle
Equity Research Analyst, Pacific Crest

Great. Listening to your presentation, clearly Oracle is leveraging both your software and your hardware business. Do you think the big data software advances you've made can accelerate the hardware side of the business?

Andrew Mendelsohn
Executive Senior Vice President, Oracle

The biggest part of our business that's here right now on the hardware side, of course, is Exadata. Exadata originally was being used almost 100% for doing big data analytics problems, big data warehousing problems, and it's still about 50% of all the Exadatas are being used in this space. We also, of course, are pretty successful with Exalytics. It's a great platform for BI tools. Our Big Data Appliance is a more recent addition. We're also doing reasonably well selling hardware into that space as well. One of the big things here that I think is worth emphasizing, once customers use these engineered systems, they're almost always buying more. It's always kicking the tires the first time around, but we see customers who are very happy with these products and become very large repeat customers.

Brendan Barnicle
Equity Research Analyst, Pacific Crest

Great. As you mentioned in your presentation, there are a lot of vendors in the big data space. They talk about different approaches to working with data than what we've seen in the past. Do you think these new approaches are going to ultimately be replacements to existing technologies or just supplements?

Andrew Mendelsohn
Executive Senior Vice President, Oracle

Yeah, I think what I'll talk about is two of the main things people are talking about in this space. There's Hadoop and there's NoSQL databases. Why don't I start with Hadoop first, since I think that is the real core of what we're talking about here in big data. The key thing to understand is what we are doing today is sort of an extension of what people have been doing for many years in the past. It's just the next generation of analytics. In the past, before Hadoop existed, what did people do? They would use file systems as staging areas for ingesting large amounts of data that they'd eventually process and bring into their data warehouse. Hadoop is really replacing that.

HDFS is a much more scalable file system. That's sort of replacing the old traditional file systems people might have been using in their analytics initiatives. Hadoop has this added benefit that not only is a good low-cost file system, good place for ingesting information, it also has a MapReduce batch processing engine for doing some analytics there. The place where people start getting confused is they think, oh, because there's some simple SQL analytics on HDFS, that suddenly they don't need massively parallel relational databases anymore. That's where they get confused. What I was trying to explain earlier is that you need the massively parallel relational databases to give interactive response times for your data scientists to do their big data analytics.

Hadoop is great, but it's really more replacing the use of file systems and ETL engines in middle-tier platforms, and it's not really a replacement for MPP relational databases. Because these relational databases are constantly raising the bar on what's normal, what's expected of them, what's table stakes, like for example, this new in-memory column store technology that we've been adding. It's not clear to me that there's any way that adding SQL on top of Hadoop is going to catch up to them anytime over the next decade or so. I think that concern is very overblown. I think the best way to look at it is Hadoop and relational databases are very complementary. NoSQL databases are an interesting area. Again, this is a technology that has been around for years and years.

It was originally called on the mainframe 30 or 40 years ago, index sequential access methods, now they're called key-value stores. These technologies, again, have been used in conjunction with relational databases for many years. They're not really big data technologies in the sense that you can do analytics against them. They don't support SQL, that's why they're called NoSQL. They're not really suitable for BI or analytics, but they are suitable for ingesting information, just like Hadoop is good at ingesting information into a file system. NoSQL databases are also good for ingestion as part of this big data story. You can get information out of them, they're, like I said, they're just key-value stores. They're not really massively parallel analytic engines. That opportunity is there, but it's sort of a smaller part of this whole big data space.

I think I'll leave it at that. There's hundreds of other vendors, I think those are the two key main ones to look at here.

Brendan Barnicle
Equity Research Analyst, Pacific Crest

One of the Hadoop distribution vendors that you guys have been working with is Cloudera, and you mentioned them in your presentation. Why did you choose Cloudera over a couple of the other distributions that are out there?

Andrew Mendelsohn
Executive Senior Vice President, Oracle

Cloudera certainly is the largest and most mature of the Hadoop distributions out there. They also have a very large and mature support organization that works with us in supporting our customers. Certainly, at the time we chose them, I think they were the clear choice, and they've chosen to be a very good partner with us as we've gone to market around big data with them.

Brendan Barnicle
Equity Research Analyst, Pacific Crest

As we step back from all this, Andy Mendelsohn, and you've looked at this over years of experience on the database side, what do you see as the biggest barriers to adoption, both for your technology and more generally for big data technology?

Andrew Mendelsohn
Executive Senior Vice President, Oracle

Well, there's these two key parts of our platform here. There's Hadoop and there's relational database technology. On the relational database technology side, I think the barriers to adoption are pretty low these days. Like I said earlier, there's a huge ecosystem of people who know how to manage Oracle databases, they know how to write SQL queries against them, and they know how to use tools that automatically generate SQL, and there's solutions, and there's whole tech stacks worth of stuff there to make it very easy. On the other side of the world, however, on the Hadoop side of the world, there are significant barriers to adoption.

I think, number one, there isn't a lot of expertise in the IT organizations on how to run Hadoop, which is why I think our engineered system for Hadoop, the Big Data Appliance, is going to really resonate with our customers there to make it easy to accept the initial deployment of the technology. The other big issue around Hadoop is, if you want to write analytics, you can start out saying, "Okay, I'm going to do some Java coding and write MapReduce using Java." Well, that is a skill set that is not very plentiful, there aren't a lot of developers out there who know how to do that. They have started building some tools on top of Java MapReduce, things like Hive and Pig, that give you a little higher-level programming paradigm. That's sort of like a simple subset of SQL. Again, that's good.

It's a good start. Most of our customers will use those kind of tools rather than try to write with Java MapReduce, that helps a little bit. I think, moving forward, they're going to, of course, working to continue that. Oracle is working with extending our SQL engine capabilities against HDFS data as well to make it easier for customers to use the Hadoop platform. Then moving up the stack, again, there's not a lot of tooling or solutions that sit on the stack today on the Hadoop side. Again, that needs to really improve to make that much more easy to adopt part of the platform. Of course, at Oracle, we'll be working on those kind of solutions as well.

Brendan Barnicle
Equity Research Analyst, Pacific Crest

The last question from me, Andy, as you look at that, maybe you sort of answered it with this previous question, where's the biggest opportunity for Oracle in the whole big data market?

Andrew Mendelsohn
Executive Senior Vice President, Oracle

Yeah. I run the database group, we see this whole big data space as being a huge market opportunity for us. Over the years, traditional BI and data warehousing has been a huge part of our business. We see big data as being just an acceleration of that business. Certainly, our Exadata engineered system is really sort of leading the way there. All the customers, I think, who have big Oracle data warehouses these days on non-Exadata platforms, as they refresh those platforms, are moving to Exadata. That's really moving forward. As I mentioned earlier, we're continuing to innovate very aggressively in this space with our new in-memory column store technology. I think that's, again, going to sort of raise the bar of what's expected of a massively parallel SQL engine.

That's going to be very hard for people in the open source space to keep pace with. I think the relational database and Exadata, of course, are huge opportunities for us. The Big Data Appliance, as customers start adopting Hadoop, is going to be another opportunity for us that's significant moving forward. In the BI space, of course, all our BI tools are moving into the in-memory analytics space. Our Exalytics engineered system is sort of the platform for doing that. That's another big opportunity for us. Then last, I did mention our Oracle NoSQL Database. We do see a lot of interest in NoSQL, not necessarily just in what we call the big data space, but just in a general data processing space that a lot of web developers especially are very interested in using NoSQL databases.

We have a very strong NoSQL offering, we're busy right now trying to make sure all of our enterprise customers know that if they are considering NoSQL, they should consider the Oracle NoSQL Database product in their evaluation. We think we're going to do very well as people do actual competitive POCs using our technology versus the other popular technologies out there for NoSQL. I'd say those are the key opportunities for us in big data.

Brendan Barnicle
Equity Research Analyst, Pacific Crest

Great. Well, that's plentiful. Andy, thanks so much for your time today. Really appreciate it. Shauna, I haven't had any questions come in. Are there any that you'd like to ask before we close things up?

Shauna O'Boyle
Senior Manager of Investor Relations, Oracle

No, I think at this point, we'll go ahead and wrap up. We'd like to thank everyone for joining us today. We'd like to extend a very special thank you to Brendan for moderating the Q&A portion of today's call and asking the questions most asked by investors. If you have any follow-up questions, please contact the investor relations team here at Oracle. This concludes our call.

Operator

Ladies and gentlemen, thank you for participating in today's conference. This does conclude the program, and you may all disconnect. Everyone, have a great day.