Tesla, Inc. (TSLA)
NASDAQ: TSLA · Real-Time Price · USD
364.27
-1.93 (-0.53%)
At close: Sep 18, 2026, 4:00 PM EDT
364.18
-0.09 (-0.02%)
After-hours: Sep 18, 2026, 7:59 PM EDT
← View all transcripts

Autonomy Investor Day 2019

Apr 22, 2019

Martin Viecha
Director of Investor Relations, Tesla

Welcome to our very first analyst day for autonomy. I really hope that this is something we can do a little more regularly now to keep you posted about the development we're doing with regards to autonomous driving. About three months ago, we were getting prepped up for our Q4 earnings call with Elon and quite a few other executives, and one of the things that I told the group is that from all the conversations that I keep having with investors on a regular basis, the biggest gap that I see with what I see inside the company and what the outside perception is our ability of autonomous driving.

It kind of makes sense because, for the past couple of years, we've been really talking about Model 3 ramp and a lot of the debate has revolved around Model 3. In reality, a lot of things have been happening in the background. We've been working on the new Full Self-Driving chip. We've had a complete overhaul of our neural net for vision recognition, et cetera.

Now that we finally started to produce our Full Self-Driving computer, we thought it's a good idea to just open the veil, invite everyone in, and talk about everything that we've been doing for the past two years. About three years ago, we wanted to find the best possible chip for full autonomy. We found out that there's no chip that's been designed from ground up for neural nets. We invited my colleague, Pete Bannon, the VP of Silicon Engineering, to design such chip for us. He's got about 35 years of experience of building chips and designing chips.

About 12 of those years were for a company called P.A. Semi, which was later acquired by Apple. He worked on dozens of different architectures and designs, and he was the lead designer, I think, for Apple iPhone 5 just before joining Tesla. He's going to be joined on the stage by Elon Musk. Thank you.

Elon Musk
CEO, Tesla

Actually, I was going to introduce Pete, but Martin's done so. Pete's just the best chip and system architect that I know in the world, and it's an honor to have you and your team at Tesla. Take it away, just tell them about the incredible work that you and your team have done.

Pete Bannon
VP of Hardware Engineering, Tesla

Thanks, Elon. It's a pleasure to be here this morning and a real treat really to tell you about all the work that my colleagues and I have been doing here at Tesla for the last three years. I think we'll tell you a little bit about how the whole thing got started. Then I'll introduce you to the Full Self-Driving computer and tell you a little bit about how it works. We'll dive into the chip itself and go through some of those details. I'll describe how the custom neural network accelerator that we designed works. Then I'll show you some results, and hopefully you'll all still be awake by then.

I was hired in February of 2016. I asked Elon if he was willing to spend all the money it takes to do full custom system design. He said, well, are we going to win? I said, well, yeah, of course. He said, I'm in. That got us started. We hired a bunch of people and started thinking about what a custom-designed chip for full autonomy would look like. We spent 18 months doing the design.

In August of 2017, we released the design for manufacturing. We got it back in December. It powered up. It actually worked very well on the first try. We made a few changes and released a B0 Rev in April of 2018. In July of 2018, the chip was qualified. We started full production of production quality parts.

In December of 2018, we had the autonomous driving stack running on the new hardware. We were able to start retrofitting employee cars and testing the hardware and software out in the real world. Just last March, we started shipping the new computer in the Model S and X. Just earlier in April, we started production in the Model 3.

This whole program, from the hiring of the first few employees to having it in full production in all three of our cars, is just a little over three years and is probably the fastest system development program I've ever been associated with. It really speaks a lot to the advantages of having a tremendous amount of vertical integration to allow you to do concurrent engineering and speed up deployment.

In terms of goals, we were totally focused exclusively on Tesla requirements. That makes life a lot easier. If you have one and only one customer, you don't have to worry about anything else. One of those goals was to keep the power under 100 W so that we could retrofit the new machine into the existing cars.

We also wanted a lower part cost so we could enable full redundancy for safety. At the time, we had a thumb-in-the-wind estimate that it would take at least 50 trillion operations a second of neural network performance to drive a car. We wanted to get at least that much, and really as much as we possibly could. Batch size is how many items you operate on at the same time.

For example, Google's TPU v1 has a batch size of 256, and you have to wait around until you have 256 things to process before you can get started. We didn't want to do that, so we designed our machine with a batch size of one, so as soon as an image shows up, we process it immediately to minimize latency, which maximizes safety. We needed a GPU to run some post-processing.

At the time, we were doing quite a lot of that, but we speculated that over time, the amount of post-processing on the GPU would decline as the neural networks got better and better. That has actually come to pass. We took a risk by putting a fairly modest GPU in the design, as you'll see, and that turned out to be a good bet. Security is super important.

If you don't have a secure car, you can't have a safe car, so there's a lot of focus on security and then, of course, safety. In terms of actually doing the chip design, as Elon alluded earlier, there was really no ground-up neural network accelerator in existence in 2016. Everybody out there was adding instructions to their CPU or GPU or DSP to make it better for inference, but nobody was really just doing it natively. We set out to do that ourselves.

For other components on the chip, we purchased industry-standard IP for CPUs and GPUs. That allowed us to minimize the design time and also the risk to the program. Another thing that was a little unexpected when I first arrived was our ability to leverage existing teams at Tesla.

Tesla had wonderful power supply design teams, signal integrity analysis, package design, system software, firmware, board designs, and a really good system validation program that we were able to take advantage of to accelerate this program. Here's what it looks like. Over there on the right, you see all the connectors for the video that comes in from the eight cameras that are in the car. You can see the two self-driving computers in the middle of the board. On the left is the power supply and some control connections.

I really love it when a solution is boiled down to its barest elements. You have video computing and power, and it's straightforward and simple. Here's the original Hardware 2.5 enclosure that the computer went into, and we've been shipping for the last two years. Here's the new design for the FSD computer.

It's basically the same. That, of course, is driven by the constraints of having a retrofit program for the cars. I'd like to point out that this is actually a pretty small computer. It fits behind the glove box, between the glove box and the firewall in the car. It does not take up half your trunk.

As I said earlier, there's two fully independent computers on the board. You can see them there highlighted in blue and green. To either side of the large SOC, you can see the DRAM chips that we use for storage, and then below left, you see the flash chips that represent the file system. These are two independent computers that boot up and run their own operating system.

Elon Musk
CEO, Tesla

Yeah, if I can add something. The general principle here is that any part of this could fail and the car will keep driving. You could have cameras fail, you could have power circuits fail, you could have one of the Tesla Full Self-Driving computer chips fail, car keeps driving. The probability of this computer failing is substantially lower than somebody losing consciousness. That's the key metric. At least an order of magnitude.

Pete Bannon
VP of Hardware Engineering, Tesla

Yep. An additional thing we do to keep the machine going is to have redundant power supplies in the car. One machine's running on one power supply and the other one's on the other. The cameras are the same. Half of the cameras run on the blue power supply, the other half run on the green power supply. Both chips receive all of the video and process it independently.

In terms of driving the car, the basic sequence is collect lots of information from the world around you. Not only do we have cameras, we also have radar, GPS, maps, the IMUs, ultrasonic sensors around the car. We have wheel ticks, steering angle. We know what the acceleration and deceleration of the car is supposed to be. All of that gets integrated together to form a plan.

Once we have a plan, the two machines exchange their independent version of the plan to make sure it's the same. Assuming that we agree, we then act and drive the car. Now, once you've driven the car with some new control, you of course want to validate it. We validate that what we transmitted was what we intend to transmit to the other actuators in the car.

Then you can use the sensor suite to make sure that it happens. If you ask the car to accelerate or brake or steer right or left, you can look at the accelerometers and make sure that you are in fact doing that. There's a tremendous amount of redundancy and overlap in both our data acquisition and our data monitoring capabilities here. Moving on to talk about the Full Self-Driving chip a little bit.

It's packaged in a 37.5 mm BGA with 1,600 balls. Most of those are used for power and ground, but plenty for signal as well. If you take the lid off, it looks like this. You can see the package substrate, and you can see the die sitting in the center there. If you take the die off and flip it over, it looks like this.

There's 13,000 C4 bumps scattered across the top of the die, and then underneath that are 12 metal layers, which is obscuring all the details of the design. If you strip that off, it looks like this. This is a 14 nm FinFET CMOS process. It's 260 mm in size, which is a modest-sized die. For comparison, a typical cell phone chip is about 100 mm 2, we're quite a bit bigger than that.

A high-end GPU would be more like 600-800 mm2 . We're in the middle. I would call it the sweet spot. It's a comfortable size to build. There's 250 million logic gates on there and a total of 6 billion transistors, which even though I work on this all the time, that's mind-boggling to me.

The chip is manufactured and tested to AEC-Q100 standards, which is a standard automotive criteria. I'd like to walk around the chip and explain all the different pieces to it, and I'm going to go in the order that a pixel coming in from the camera would visit all the different pieces. Up there on the top left, you can see the camera serial interface.

We can ingest 2.5 billion pixels per second, which is more than enough to cover all the sensors that we know about. We have an on-chip network that distributes data from the memory system. The pixels would travel across the network to the memory controllers on the right and left edges of the chip.

We use industry-standard LPDDR4 memory running at 4,266 Gb per second, which gives us a peak bandwidth of 68 GB a second, which is a pretty healthy bandwidth. Again, this is not ridiculous. We're trying to stay in the comfortable sweet spot for cost reasons. The image signal processor has a 24-bit internal pipeline that allows us to take full advantage of the HDR sensors that we have around the car.

It does advanced tone mapping, which helps to bring out details and shadows. It has advanced noise reduction, which improves the overall quality of the images that we're using in the neural network. The neural network accelerator itself, there's two of them on the chip. They each have 32 MB of SRAM to hold temporary results and minimize the amount of data that we have to transmit on and off the chip, which helps reduce power.

Each array has a 96 by 96 multiply add array with in-place accumulation, which allows us to do almost 10,000 multiply adds per cycle. There's dedicated ReLU hardware, dedicated pooling hardware, and each of these deliver 306. Excuse me. Each one delivers 36 trillion operations per second, and they operate at 2 GHz. The two of them together on a die deliver 72 trillion operations a second.

We exceeded our goal of 50 TOPS by a fair bit. There's also a video encoder. We encode video and use it in a variety of places in the car, including the backup camera display. There's optionally a user feature for a dash cam and also for clip logging data to the cloud, which Stuart and Andrej will talk about more later.

There's a GPU on the chip. It's modest performance. It has support for both 32 and 16-bit floating point. We have 12 A72 64-bit CPUs for general purpose processing. They operate at 2.2 GHz, and this represents about 2.5 x the performance available in the current solution. There's a safety system that contains two CPUs that operate in lockstep. This system is the final arbiter of whether it's safe to actually drive the actuators in the car.

This is where the two plans come together, and we decide whether it's safe or not to move forward. Lastly, there is a safety system. Basically, the job of the safety system is to ensure that this chip only runs software that's been cryptographically signed by Tesla. If it's not been signed by Tesla, the chip does not operate.

I've told you a lot of different performance numbers, I thought it'd be helpful maybe to put it into perspective a little bit. Throughout this talk, I'm going to talk about a neural network from our narrow camera that uses 35 billion operations, 35 GOPS. If we use all 12 CPUs to process that network, we could do one and a half frames per second, which is super slow and not nearly adequate to drive the car.

If we use the 600 Gflop GPU, the same network, we'd get 17 frames per second, which is still not good enough to drive a car with eight cameras. The neural network accelerators on the chip can deliver 2,100 frames per second. You can see from the scaling as we moved along, that the amount of computing in the CPU and GPU are basically insignificant to what's available in the neural network accelerator. It really is night and day.

Moving on to talk about the neural network accelerator. I'm just going to stop for some water. On the left, there is a cartoon of a neural network, just to give you an idea of what's going on. The data comes in at the top and visits each of the boxes, and the data flows along the arrows to the different boxes.

The boxes are typically convolutions or deconvolution with ReLUs. The green boxes are pooling layers. The important thing about this is that the data produced by one box is then consumed by the next box, and then you don't need it anymore. You can throw it away. All of that temporary data that gets created and destroyed as you flow through the network, there's no need to store that off-chip in DRAM.

We keep all that data in SRAM, and I'll explain why that's super important in a few minutes. If you look over on the right side of this, you can see that in this network of the 35 billion operations, almost all of them are convolution, which is based on dot products. The rest are deconvolution, also based on dot product, and then ReLU and pooling, which are relatively simple operations.

If you were designing some hardware, you'd clearly target doing dot products, which are based on multiply add and really kill that. Imagine that you sped it up by a factor of 10,000. 100% all of a sudden turns into 0.1%, 0.01%, and suddenly the ReLU and pooling operations are going to be quite significant. Our hardware design includes dedicated resources for processing ReLU and pooling as well.

This chip is operating in a thermally constrained environment, so we had to be very careful about how we burn that power. We want to maximize the amount of arithmetic we can do. We picked integer add. It's 9x less energy than a corresponding floating point add. We picked eight-bit by eight-bit integer multiply, which is significantly less power than other multiply operations and is probably enough accuracy to get good results.

In terms of memory, we chose to use SRAM as much as possible, you can see there that going off-chip to DRAM is approximately 100 x more expensive in terms of energy consumption than using local SRAM. Clearly, we want to use local SRAM as much as possible.

In terms of control, this is data that was published in a paper by Mark Horowitz at ISSCC, where he sort of critiqued how much power it takes to execute a single instruction on a regular integer CPU. You can see that the add operation is only 0.15% of the total power. All the rest of the power is control overhead and bookkeeping. In our design, we sought to basically get rid of all that as much as possible because what we're really interested in is arithmetic.

Here's the design that we finished. You can see that it's dominated by the 32 MB of SRAM. There's big banks on the left and right and in the center bottom. All the computing is done in the upper middle. Every single clock, we read 256 B of activation data out of the SRAM array, 128 B of weight data out of the SRAM array, and we combine it in a 96 by 96 muladd array, which performs 9,000 multiply adds per clock.

At 2 GHz, that's a total of 36.8 TOPS. Now, when we're done with the dot product, we unload the engine so that we shift the data out across the dedicated ReLU unit, optionally across a pooling unit, and then finally into a write buffer where all the results get aggregated up, and then we write out 128 B per cycle back into the SRAM.

This whole thing cycles along all the time continuously, we're doing dot products while we're unloading previous results, doing pooling, and writing back into the memory. If you add it all up, at 2 GHz, you need one terabyte per second of SRAM bandwidth to support all that work, the hardware supplies that. 1 TB per second of bandwidth per engine. There's two on the chip, 2 TB per second.

The accelerator has a relatively small instruction set. We have a DMA read operation to bring data in from memory. We have a DMA write operation to push results back out to memory. We have three dot product-based instructions, convolution, deconvolution, and inner product. Then two relatively simple, scale is a one input, one output operation, and outwise is two inputs and one output. Of course, stop when you're done.

We had to develop a neural network compiler for this. We take the neural network that's been trained by our vision team as it would be deployed in the older cars, we take that and compile it for use on the new accelerator. The compiler does layer fusion, which allows us to maximize the computing each time we read data out of the SRAM and put it back.

It also does some smoothing so that the demands on the memory system aren't too lumpy. We also do channel padding to reduce bank conflicts, and we do bank-aware SRAM allocation. This is a case where we could've put more hardware in the design to handle bank conflicts, but by pushing it into software, we save hardware and power at the cost of some software complexity.

We also automatically insert DMAs into the graph so that data arrives just in time for computing without having to stall the machine. At the end, we generate all the code, we generate all the weight data, we compress it, and we add a CRC checksum for reliability. To run a program, all the neural network descriptions programs are loaded into SRAM, at the start, and then they sit there ready to go all the time.

To run a network, you have to program the address of the input buffer, which presumably is a new image that just arrived from a camera. You set the output buffer address, you set the pointer to the network weights, and then you set go. The machine goes off and will sequence through the entire neural network all by itself, usually running for 1 million or 2 million cycles. When it's done, you get an interrupt and can post-process the results.

Moving on to results. We had a goal to stay under 100 W. This is measured data from cars driving around running the full Autopilot stack. We're dissipating 72 W, which is a little bit more power than the previous design, but with the dramatic improvement in performance, it's still a pretty good answer.

Of that 72 W, about 15 W is being consumed running the neural networks. In terms of cost, the silicon cost of this solution is about 80% of what we were paying before. We are saving money by switching to this solution. In terms of performance, we took the narrow camera neural network, which I've been talking about, that has 35 billion operations in it.

We ran it on the old hardware in a loop as quick as possible, and we delivered 110 frames per second. We took the same data, the same network, compiled it for hardware for the new FSD computer. Using all four accelerators, we can get 2,300 frames per second processed. A factor of 21.

Elon Musk
CEO, Tesla

I think this is perhaps the most significant slide. It's night and day.

Pete Bannon
VP of Hardware Engineering, Tesla

I've never worked on a project where the performance increase was more than three. This was pretty fun. If you compare it to, say, NVIDIA's DRIVE Xavier solution, a single chip delivers 21 TOPS. Our Full Self-Driving computer with two chips is 144 TOPS. To conclude, I think we've created a design that delivers outstanding performance, 144 TOPS for neural network processing.

It has outstanding power performance. We managed to jam all of that performance into the thermal budget that we had. It enables a fully redundant computing solution. It has a modest cost. Really, the important thing is that this FSD computer will enable a new level of safety and autonomy in Tesla's vehicles without impacting their cost or range. Something that I think we're all looking forward to.

Elon Musk
CEO, Tesla

I think why don't we do Q&A after each segment, so if people have questions about the hardware, they can ask right now. The reason I asked Pete to do just a detailed, far more detailed than perhaps most people would appreciate a dive into the Tesla Full Self-Driving computer is because at first it seems improbable. How could it be that Tesla, who has never designed a chip before, would design the best chip in the world?

That is objectively what has occurred. Not best by a small margin, best by a huge margin. It's in the cars right now. All Teslas being produced right now have this computer. We switched over from the NVIDIA solution for S and X about a month ago, and we switched over Model 3 about 10 days ago. All cars being produced have all the hardware necessary, compute and otherwise, for Full Self-Driving.

I'll say that again. All Tesla cars being produced right now have everything necessary for Full Self-Driving. All you need to do is improve the software. Later today, you will drive the cars with the development version of the improved software, and you will see for yourselves. Questions for Pete?

Pete Bannon
VP of Hardware Engineering, Tesla

Yep.

Trip Chowdhry
Analyst, Global Equities Research

Very impressive in every shape and form. Questions, I saw some of those Trip Chowdhry, Global Equities Research. Very impressive in every shape and form. I was wondering, I took some notes. You are using activation function, ReLU, the rectified linear unit. If we think about the deep neural network, it has multiple layers, and some algorithms may use different activation functions for different hidden layers, like Softmax or TANH. Do you have flexibility for incorporating different activation functions rather than ReLU in your platform? I have a follow-up.

Pete Bannon
VP of Hardware Engineering, Tesla

Yes, we do. We have implementations of TANH and Sigmoid, for example.

Trip Chowdhry
Analyst, Global Equities Research

Beautiful. One last question. In the nanometers, you mentioned 14 nm. I was wondering, wouldn't it make sense to come a little lower, maybe 10 nm two years down or maybe seven?

Pete Bannon
VP of Hardware Engineering, Tesla

At the time we started the design, not all the IP that we wanted to purchase was available in 10 nm, so we finished the design in 14 nm.

Elon Musk
CEO, Tesla

It's maybe worth pointing out that we finished this design maybe one and a half, two years ago, and began design of the next generation. We're not talking about the next generation today, but we're about halfway through it. All the things that are obvious for a next-generation chip, we're doing.

Pete Bannon
VP of Hardware Engineering, Tesla

Yeah. Yep. Oh. Oh, hi.

Adam Jonas
Analyst, Morgan Stanley

You talked about the software as the piece now. You did a great job. I was blown away. Understood 10% of what you said, but I trust that it's in good hands.

Pete Bannon
VP of Hardware Engineering, Tesla

Thanks.

Adam Jonas
Analyst, Morgan Stanley

It feels like you got the hardware piece is done, and that was really hard to do, and now you have to do the software piece. Maybe that's outside of your expertise. How should we think about that software piece?

Elon Musk
CEO, Tesla

Well, couldn't ask for a better introduction to Andrej and Stuart.

Pete Bannon
VP of Hardware Engineering, Tesla

I think-

Elon Musk
CEO, Tesla

Are there any questions for the chip part? The next part of the presentation is neural nets and software.

Adam Jonas
Analyst, Morgan Stanley

Let me-

Elon Musk
CEO, Tesla

Yeah.

Adam Jonas
Analyst, Morgan Stanley

Maybe on the chip side, the last slide was 144 trillions of operations per second versus, was it NVIDIA at 21?

Pete Bannon
VP of Hardware Engineering, Tesla

That's right.

Adam Jonas
Analyst, Morgan Stanley

Maybe can you just contextualize that for a finance person, why that's so significant, that gap? Thank you.

Pete Bannon
VP of Hardware Engineering, Tesla

Well, it's a factor of seven in performance delta, that means you can do seven times as many frames. You can run neural networks that are seven times larger and more sophisticated. It's a very big currency that you can spend on lots of interesting things to make the car better.

Elon Musk
CEO, Tesla

I think the Xavier power usage is higher than ours. Xavier power is higher than ours, I think. Or comparable.

Pete Bannon
VP of Hardware Engineering, Tesla

I don't know that.

Elon Musk
CEO, Tesla

To the best of my knowledge, the power requirements would increase at least to the same degree, a factor of seven, and costs would also increase by a factor of seven. Okay, yeah.

Pete Bannon
VP of Hardware Engineering, Tesla

Great. Shall we?

Elon Musk
CEO, Tesla

Power is a real problem because it also reduces range. The penalty for power is very high. Then you have to get rid of that power. The thermal problem becomes really significant because you've got to get rid of all that power.

Adam Jonas
Analyst, Morgan Stanley

Thank you very much. I think we have quite a bit.

Elon Musk
CEO, Tesla

Just ask the questions. If you guys don't mind the day running a bit long, we're going to do the drive demos afterwards. If anybody needs to pop out and do drive demos a little sooner, you're welcome to do that. I do want to make sure we answer your questions.

Pete Bannon
VP of Hardware Engineering, Tesla

Yep.

Pradeep Ramani
Analyst, UBS

Pradeep Ramani from UBS. Intel and AMD, to some extent, have started moving towards a chiplet-based architecture. I did not notice a chiplet-based design here. Do you think that looking forward, that would be something that might be of interest to you guys from an architecture standpoint?

Pete Bannon
VP of Hardware Engineering, Tesla

A chiplet-based architecture?

Pradeep Ramani
Analyst, UBS

Yes.

Pete Bannon
VP of Hardware Engineering, Tesla

We're not currently considering anything like that. I think that's mostly useful when you need to use different styles of technology. If you want to integrate silicon germanium or DRAM technology on the same silicon substrate, that gets pretty interesting. Until the die size gets obnoxious, I wouldn't go there.

Elon Musk
CEO, Tesla

Okay. To be clear, the strategy here, and this started basically a little over three years ago, was design and build a computer that is fully optimized and aiming for Full Self-Driving, write software that is designed to work specifically on that computer and get the most out of that computer. You have tailored hardware that is a master of one trade, self-driving.

NVIDIA is a great company, they have many customers, as they apply their resources, they need to do a generalized solution. We care about one thing, self-driving. It was designed to do that incredibly well. The software is also designed to run on that hardware incredibly well. The combination of the software and the hardware, I think, is unbeatable.

Pradeep Ramani
Analyst, UBS

Sure.

Speaker 16

Hi. The chip is designed to process video input. In case you use, let's say, LiDAR, would it be able to process that as well, or is it primarily for video input?

Pete Bannon
VP of Hardware Engineering, Tesla

To be honest-

Elon Musk
CEO, Tesla

What we're going to explain to you today is that LiDAR is a fool's errand, and anyone relying on LiDAR is doomed. Doomed. Expensive sensors that are unnecessary. It's like having a whole bunch of expensive appendices. One appendix is bad. Now we're going to put a whole bunch of them. That's ridiculous. You'll see.

Pete Bannon
VP of Hardware Engineering, Tesla

There's somebody up here.

Speaker 16

Hi.

Pete Bannon
VP of Hardware Engineering, Tesla

Oh, there's a gentleman with a mic. Hi.

Speaker 16

Hi. Just two questions. Just on the power consumption, is there a way to maybe give us a rule of thumb on every watt reduces range by a certain percent or a certain amount, just so we can get a sense of how much of-

Pete Bannon
VP of Hardware Engineering, Tesla

Well, a Model 3, the target consumption is 250 W per mile.

Elon Musk
CEO, Tesla

It depends on the nature of the driving as to how many miles that affects. In city, it would have a much bigger effect than on highway. If you're driving for an hour in a city, and you had a solution, hypothetically, that was a kilowatt, you'd lose four miles on a Model 3. If you're only going, say, 12 mi an hour, then that would be a 25% impact on range in city. Basically, the power of the system has a massive impact on city range, which is where we think most of the Robotaxi market will be. Power is extremely important.

Pete Bannon
VP of Hardware Engineering, Tesla

Okay. Tasha? I'm sorry, I didn't hear you.

Tasha Keeney
Analyst, ARK Invest

Sorry. Thank you. What's the primary design objective of the next generation chip?

Elon Musk
CEO, Tesla

We don't want to talk too much about the next generation chip but it-

Pete Bannon
VP of Hardware Engineering, Tesla

Safety.

Elon Musk
CEO, Tesla

... it'll be at least, let's say, 3x better than the current system.

Pete Bannon
VP of Hardware Engineering, Tesla

Mike, go ahead.

Elon Musk
CEO, Tesla

It's about two years away.

Speaker 17

You mentioned that about cost being less to develop this chip. You don't manufacture the chip. You contract that out. How much cost reduction does that save in the overall vehicle cost?

Pete Bannon
VP of Hardware Engineering, Tesla

The 20% cost reduction I cited was the piece cost per vehicle reduction. That wasn't a development cost, that was just the actual.

Speaker 17

No, I'm saying, if I'm manufacturing these en masse, is this saving money in doing it yourself?

Pete Bannon
VP of Hardware Engineering, Tesla

Yes, a little bit.

Elon Musk
CEO, Tesla

Most people don't make chips with their own fab. It's pretty unusual. Yeah.

Speaker 17

You don't see any supply issues-

Elon Musk
CEO, Tesla

No.

Speaker 17

... getting the chip mass-produced?

Elon Musk
CEO, Tesla

No.

Pete Bannon
VP of Hardware Engineering, Tesla

The cost saving pays for the development. The basic strategy going to Elon was, we're going to build this chip. It's going to reduce the cost. Elon said, hmm, times a million cars a year. Deal.

Speaker 17

No more digging for this model?

Elon Musk
CEO, Tesla

That's correct.

Pete Bannon
VP of Hardware Engineering, Tesla

Yes. Sorry.

Elon Musk
CEO, Tesla

If they're really chip-specific questions, we can answer them. Otherwise, there will be a Q&A opportunity after-

Pete Bannon
VP of Hardware Engineering, Tesla

Yeah.

Elon Musk
CEO, Tesla

... Andrej talks and after Stuart talks. There will be two other Q&A opportunities. If it's very chip specific, then-

Pete Bannon
VP of Hardware Engineering, Tesla

I'll be here all afternoon.

Elon Musk
CEO, Tesla

I think people will be here at the end as well. Go ahead.

Ben Kallo
Analyst, Baird

Hi. Thanks. That die photo you had, the neural processor takes up quite a bit of the die. I'm curious, is that your own design, or is there some external IP there?

Pete Bannon
VP of Hardware Engineering, Tesla

Yes. That was a custom design by Tesla.

Ben Kallo
Analyst, Baird

Okay. I guess the follow-on would be, there's probably a fair amount of opportunity to reduce that footprint as you tweak the design?

Pete Bannon
VP of Hardware Engineering, Tesla

It's actually quite dense. In terms of reducing it, I don't think so. It'll greatly enhance the functional capabilities in the next generation.

Ben Kallo
Analyst, Baird

Okay. Last question, can you share where you're fabbing this part?

Elon Musk
CEO, Tesla

Wait, what?

Pete Bannon
VP of Hardware Engineering, Tesla

Where are we fabbing it?

Elon Musk
CEO, Tesla

Oh, Samsung.

Ben Kallo
Analyst, Baird

Samsung?

Pete Bannon
VP of Hardware Engineering, Tesla

Yes.

Ben Kallo
Analyst, Baird

All right.

Pete Bannon
VP of Hardware Engineering, Tesla

Austin, Texas.

Ben Kallo
Analyst, Baird

Thank you.

Pete Bannon
VP of Hardware Engineering, Tesla

There's one at the back.

Grant Tanaka
Analyst, Grant Tanaka Capital Management

Hi. Grant Tanaka Capital Management. Just curious how defensible your chip technologies and design is from an IP point of view, and hoping that you won't be offering a lot of the IP to the outside for free. Thanks.

Pete Bannon
VP of Hardware Engineering, Tesla

We have filed on the order of a dozen patents on this technology. Fundamentally, it's linear algebra, which I don't think you can patent. I'm not sure, we'll try.

Elon Musk
CEO, Tesla

I think if somebody started today and they were really good, they might have something like what we have right now in three years. In two years, we'll have something 3x better.

Trip Chowdhry
Analyst, Global Equities Research

Talking about the intellectual property protection, you have the best intellectual property, and some people just steal it for the fun of it. I was wondering, if we look at few interaction with Aurora, that industry believes they stole your intellectual property. I think the key ingredient that you need to protect is the weights that you associate to various parameters.

Do you think your chip can do something to prevent anybody, maybe encrypt all the weights so that even you don't know what the weights are at the chip level, so that your intellectual property remains inside it and nobody knows about it, and nobody can just steal it?

Elon Musk
CEO, Tesla

Man, I'd like to meet the person that could do that, because I would hire them in a heartbeat. Yeah. That'd be a hard problem. Yeah. We do encrypt. It's a hard chip to crack. If they can crack it's very good. If they can then crack it and then also figure out the software and the neural net system and everything else, they can design it from scratch. That's a lot harder.

Pete Bannon
VP of Hardware Engineering, Tesla

It's our intention to prevent people from stealing all that stuff, and if they do, we hope it at least takes a long time.

Elon Musk
CEO, Tesla

It will definitely take them a long time. Yeah. If it was our goal to do that, how would we do it? It'd be very difficult. The thing that's, I think, a very powerful, sustainable advantage for us is the fleet. Nobody has the fleet. Those weights are constantly being updated and improved based on billions of miles driven.

Tesla has 100 x more cars with the Full Self-Driving hardware than everyone else combined. By the end of this quarter, we'll have 500,000 cars with the full eight-camera setup, 12 ultrasonics. Some of them will still be on Hardware 2, but we still have the data-gathering ability. By a year from now, we'll have over a million cars with Full Self-Driving computer hardware, everything. Yeah.

Pete Bannon
VP of Hardware Engineering, Tesla

We have Andrej.

Elon Musk
CEO, Tesla

It's just a massive data advantage. It's similar to how the Google search engine has a massive advantage because people use it, and the people effectively program Google with their queries and the results.

Pete Bannon
VP of Hardware Engineering, Tesla

James?

Speaker 18

Sir, may I just press you on that? Please reframe the question because I'm a tech layman, if it's appropriate. When we talk to Waymo or NVIDIA, they do speak with equivalent conviction about their leadership because of their competence in simulating miles driven. Can you talk about the advantage of having real-world miles versus simulated miles?

I think they express that by the time you get a million miles, they can simulate a billion, and no Formula One race car driver, for example, could ever successfully complete a real-world track without driving in a simulator. Can you talk about the advantages, it sounds like that you perceive to have associated with having data ingestion coming from real-world miles versus simulated miles?

Elon Musk
CEO, Tesla

Absolutely. The simulator, we have quite a good simulation too. It just does not capture the long tail of weird things that happen in the real world. If the simulation fully captured the real world, well, that would be proof that we're living in a simulation, I think. Yeah. It doesn't. I wish. Simulations do not capture the real world. The real world's really weird and messy. You need the cars on the road. We're actually going to get into that in Andrej and Stuart's presentation.

Pete Bannon
VP of Hardware Engineering, Tesla

Yeah. Sure.

Elon Musk
CEO, Tesla

Okay, why don't we move on to Andrej?

Pete Bannon
VP of Hardware Engineering, Tesla

Great. Thanks.

Martin Viecha
Director of Investor Relations, Tesla

Thank you.

Pete Bannon
VP of Hardware Engineering, Tesla

Thank you, everybody.

Martin Viecha
Director of Investor Relations, Tesla

Thank you very much. The last question was actually a very good segue. One thing to remember about our FSD computer is that it can run much more complex neural nets for much more precise image recognition. To talk to you about how we actually get that image data and how we analyze them, we have our Senior Director of AI, Andrej Karpathy, who's going to explain all of that to you. Andrej has a PhD from Stanford University, where he studied computer science, focusing on- Actually v ision recognition and deep learning.

Elon Musk
CEO, Tesla

Andrej, why don't you just talk, do your own intro?

Martin Viecha
Director of Investor Relations, Tesla

Yeah. Exactly.

Elon Musk
CEO, Tesla

There's a lot of PhDs from Stanford, that's not important.

Martin Viecha
Director of Investor Relations, Tesla

Yes.

Elon Musk
CEO, Tesla

Okay. We don't care.

Martin Viecha
Director of Investor Relations, Tesla

Come on up. Thank you.

Elon Musk
CEO, Tesla

Andrej studied the computer vision class at Stanford. That's much more significant. That's what matters.

Martin Viecha
Director of Investor Relations, Tesla

Yeah.

Elon Musk
CEO, Tesla

Can you please talk about your background in a way that is not bashful. Just feel free to tell me about the stuff you've done.

Andrej Karpathy
Senior Director of AI, Tesla

Sure.

Elon Musk
CEO, Tesla

Yeah.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah, I think I've been training neural networks basically for what is now a decade. These neural networks were not actually really used in the industry until maybe five or six years ago. It's been some time that I've been training these neural networks, and that included institutions at Stanford, at OpenAI, at Google, and really just training a lot of neural networks, not just for images, but also for natural language and designing architectures that couple those two modalities for my PhD.

Elon Musk
CEO, Tesla

In a computer science class.

Andrej Karpathy
Senior Director of AI, Tesla

Oh, yeah. At Stanford, I actually taught the convolutional neural networks class. I was the primary instructor for that class. I actually started the course and designed the entire curriculum. In the beginning, it was about 150 students, and then it grew to 700 students over the next two or three years. It's a very popular class. It's one of the largest classes at Stanford right now.

Elon Musk
CEO, Tesla

Yeah.

Andrej Karpathy
Senior Director of AI, Tesla

That was also really successful.

Elon Musk
CEO, Tesla

Andrej is really one of the best computer vision people in the world, arguably the best.

Andrej Karpathy
Senior Director of AI, Tesla

Okay. Thank you. Yeah. Hello, everyone. Pete told you all about the chip that we've designed that runs neural networks in the car. My team is responsible for training of these neural networks, and that includes all of data collection from the fleet, neural network training, and then some of the deployment onto that chip. What do the neural networks do exactly in the car?

What we are seeing here is a stream of videos from across the vehicle, across the car. These are eight cameras that send us videos, and then these neural networks are looking at those videos and are processing them and making predictions about what they're seeing. Some of the things that we're interested in, and some of the things you're seeing on this visualization here, are lane line markings, other objects, the distances to those objects, what we call drivable space, shown in blue, which is where the car is allowed to go, and a lot of other predictions like traffic lights, traffic signs, and so on.

Now, for my talk, I will talk roughly in three stages. First, I'm going to give you a short primer on neural networks and how they work and how they're trained. I need to do this because I need to explain in the second part why it is such a big deal that we have the fleet and why it's so important, and why it's a key enabling factor to really training these neural networks and making them work effectively on the roads.

In the third stage, I'll talk about vision and LiDAR and how we can estimate depth just from vision alone. The core problem that these networks are solving in the car is that of visual recognition. For you and I, this is a very simple problem. You can look at all of these four images, and you can see that they contain a cello, a boat, an iguana, or scissors.

This is very simple and effortless for us. This is not the case for computers. The reason for that is that these images are to a computer, really just a massive grid of pixels. At each pixel, you have the brightness value at that point. Instead of just seeing an image, a computer really gets a million numbers in a grid that tell you the brightness values at all the positions.

Elon Musk
CEO, Tesla

The Matrix, if you will. It really is The Matrix.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah.

Elon Musk
CEO, Tesla

Yeah.

Andrej Karpathy
Senior Director of AI, Tesla

We have to go from that grid of pixels and brightness values into high-level concepts like iguana and so on. As you might imagine, this iguana has a certain pattern of brightness values, but iguanas actually can take on many appearances. They can be in many different appearances, different poses, and different brightness conditions against different backgrounds.

You can have different crops of that iguana. We have to be robust across all those conditions, and we have to understand that all those different brightness patterns actually correspond to iguanas. The reason you and I are very good at this is because we have a massive neural network inside our heads that is processing those images.

Light hits the retina, travels to the back of your brain to the visual cortex. The visual cortex consists of many neurons that are wired together and that are doing all the pattern recognition on top of those images. Really over the last, I would say about five years, the state-of-the-art approaches to processing images using computers have also started to use neural networks, but in this case, artificial neural networks.

These artificial neural networks, and this is just a cartoon diagram of it, are a very rough mathematical approximation to your visual cortex, really do have neurons, and they are connected together. Here I'm only showing three or four neurons in four layers, but a typical neural network will have tens to hundreds of millions of neurons, and each neuron will have 1,000 connections. These are really large pieces of almost simulated tissue.

Then what we can do is we can take those neural networks, and we can show them images. For example, I can feed my iguana into this neural network, and the network will make predictions about what it's seeing. In the beginning, these neural networks are initialized completely randomly.

The connection strengths between all those different neurons are completely random, and therefore, the predictions of that network are also going to be completely random. It might think that you're actually looking at a boat right now, and it's very unlikely that this is actually an iguana. During a training process, really what we're doing is we know that that's actually an iguana. We have a label.

What we're doing is we're basically saying we'd like the probability of iguana to be larger for this image and the probability of all the other things to go down. Then there's a mathematical process called backpropagation and stochastic gradient descent that allows us to backpropagate that signal through those connections and update every one of those connections, sorry. Update every one of those connections just a little amount, and once the update is complete, the probability of iguana for this image will go up a little bit. It might become 14%, and the probability of the other things will go down.

Of course, we don't just do this for this single image. We actually have entire large data sets that are labeled. We have lots of images. Typically, you might have millions of images, thousands of labels, or something like that, and you are doing forward-backward passes over and over again. You're showing the computer, here's an image. It has an opinion, and then you're saying, this is the correct answer, and it tunes itself a little bit. You repeat this millions of times, and sometimes you show images, the same image to the computer hundreds of times as well.

The network training typically will take on the order of a few hours or a few days, depending on how big of a network you're training. That's the process of training a neural network. There's something very unintuitive about the way neural networks work that I have to really get into, and that is that they really do require a lot of these examples, and they really do start from scratch.

They know nothing, and it's really hard to wrap your head around this. As an example, here's a cute dog, and you probably may not know the breed of this dog, but the correct answer is that this is a Japanese spaniel. All of us are looking at this, and we're seeing Japanese spaniel, and we're like, okay, I got it. I understand kind of what this Japanese spaniel looks like.

If I show you a few more images of other dogs, you can probably pick out other Japanese spaniels here. In particular, those three look like a Japanese spaniel, and the other ones do not. You can do this very quickly, and you need one example, but computers do not work like this. They actually need a ton of data of Japanese spaniels.

This is a grid of Japanese spaniels showing them, you need thousands of examples, showing them in different poses, different brightness conditions, different backgrounds, different crops. You really need to teach the computer from all the different angles what this Japanese spaniel looks like, and it really requires all that data to get that to work, otherwise, the computer can't pick up on that pattern automatically.

What does all this imply about the setting of self-driving? Of course, we don't care about dog breeds too much. Maybe we will at some point. For now, we really care about lane line markings, objects, where they are, where we can drive, and so on. The way we do this is we don't have labels like iguana for images, but we do have images from the fleet like this, and we're interested in, for example, lane line markings.

A human typically goes into an image and using a mouse annotates the lane line markings. Here's an example of an annotation that a human could create a label for this image, and it's saying that that's what you should be seeing in this image. These are the lane line markings.

Then what we can do is we can go to the fleet, and we can ask for more images from the fleet, and if you ask the fleet, if you just do a naive job of this, and you just ask for images at random, the fleet might respond with images like this, typically going forward on some highway. You might just get a random collection like this, and we would annotate all that data.

If you're not careful, and you only annotate a random distribution of this data, your network will kind of pick up on this random distribution of data and work only in that regime. If you show it a slightly different example, for example, here is an image that actually the road is curving, and it is a bit of a more residential neighborhood.

If you show the neural network this image, that network might make a prediction that is incorrect. It might say that, okay. Well, I've seen lots of times on highways, lanes just go forward, so here's a possible prediction. Of course, this is very incorrect. The neural network really can't be blamed. It does not know that the tree on the left, whether or not it matters or not.

It does not know if the car on the right matters or not towards the lane line. It does not know that the buildings in the background matter or not. It really starts completely from scratch. You and I know that the truth is that none of those things matter. What actually matters is that there are a few white lane line markings over there in the vanishing point, and the fact that they curl a little bit should pull the prediction.

Except there's no mechanism by which we can just tell the neural network, hey, those lane line markings actually matter. The only tool in the toolbox that we have is labeled data. What we do is we need to take images like this when the network fails, and we need to label them correctly.

In this case, we will turn the lane to the right, and then we need to feed lots of images of this to the neural net. A neural net over time will basically pick up on this pattern that those things there don't matter, but those lane line markings do, and will learn to predict the correct lane. What's really critical is not just the scale of the dataset.

We don't just want millions of images. We actually need to do a really good job of covering the possible space of things that the car might encounter on the roads. We need to teach the computer how to handle scenarios where it's night and wet. You have all these different specular reflections, and as you might imagine, the brightness patterns in these images will look very different.

We have to teach the computer how to deal with shadows, how to deal with forks in the road, how to deal with large objects that might be taking up most of that image, how to deal with tunnels, or how to deal with construction sites. In all these cases, there's no, again, explicit mechanism to tell the network what to do.

We only have massive amounts of data. We want to source all those images, and we want to annotate the correct lines, and the network will pick up on the patterns of those. Large and varied datasets basically make these networks work very well. This is not just a finding for us here at Tesla. This is a ubiquitous finding across the entire industry.

Experiments and research from Google, from Facebook, from Baidu, from Alphabet's DeepMind all show similar plots where neural networks really love data and love scale and variety. As you add more data, these neural networks start to work better and get higher accuracies for free. More data just makes them work better.

A number of people have kind of pointed out that potentially we could use simulation to actually achieve the scale of the datasets, and we're in charge of a lot of the conditions here, and maybe we can achieve some variety in a simulator. At Tesla, and that was also kind of brought up in the questions just before this. At Tesla, this is actually a screenshot of our own simulator. We use simulation extensively. We use it to develop and evaluate the software.

We've also even used it for training quite successfully. Really, when it comes to training data for neural networks, there really is no substitute for real data. The simulations have a lot of trouble with modeling appearance, physics, and the behaviors of all the agents around you. Here are some examples to really drive that point across. The real world really throws a lot of crazy stuff at you. In this case, for example, we have very complicated environments with snow, with trees, with wind.

We have various visual artifacts that are hard to simulate, potentially. We have complicated construction sites, bushes and plastic bags that can kind of go around with the wind. Complicated construction sites that might feature lots of people, kids, animals, all mixed in, and simulating how those things interact and flow through this construction zone might actually be completely intractable.

It's not about the movement of any one pedestrian in there. It's about how they respond to each other and how those cars respond to each other, and how they respond to you driving in that setting. All of those are actually really tricky to simulate. It's almost like you have to solve the self-driving problem to just simulate other cars in your simulation. It's really complicated. We have dogs, exotic animals. In some cases, it's not even that you can't simulate it, is that you can't even come up with it.

Elon Musk
CEO, Tesla

Yeah.

Andrej Karpathy
Senior Director of AI, Tesla

For example, I didn't know that you can have truck on truck on truck like that, in the real world, you find this, and you find lots of other things that are very hard to, really even come up with. Really, the variety that I'm seeing in the data coming from the fleet is just crazy with respect to what we have in the simulator. We have a really good simulator.

Elon Musk
CEO, Tesla

Yeah. I think simulation, you're fundamentally grading your own homework. If you know that you're going to simulate it, okay, you can definitely solve for it. As Andrej is saying, you don't know what you don't know. The world is very weird and has millions of corner cases. If somebody can produce a self-driving simulation that accurately matches reality, that in itself would be a monumental achievement of human capability. They can't. There's no way.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah. I think the three points that I really tried to drive home until now are to get neural networks to work well, you require these three essentials. You require a large dataset, a varied dataset, and a real dataset.

If you have those capabilities, you can actually train neural networks and make them work very well. Why is Tesla in such a unique and interesting position to really get all these three essentials right? The answer to that, of course, is the fleet. We can really source data from it and make our neural network systems work extremely well.

Let me take you through a concrete example of, for example, making the object detector work better to give you a sense of how we develop these neural networks, how we iterate on them, and how we actually get them to work over time. Object detection is something we care a lot about. We like to put bounding boxes around, let's say, the cars and the objects here because we need to track them, and we need to understand how they might move around.

Again, we might ask human annotators to give us some annotations for these, and humans might go in and might tell you that, okay, those patterns over there are cars and bicycles and so on. You can train your neural network on this, but if you're not careful, the neural network will make mispredictions in some cases.

As an example, if we stumble by a car like this that has a bike on the back of it, then the neural network actually, when I joined, would actually create two detections. It would create a car detection and a bicycle detection. That's actually kind of correct because I guess both of those objects actually exist. For the purposes of the controller and the planner downstream, you really don't want to deal with the fact that this bicycle can go with the car.

The truth is that bike is attached to that car, so in terms of just objects on the road, there's a single object, a single car. What you'd like to do now is you'd like to just potentially annotate lots of those images as this is just a single car.

The process that we go through internally in the team is that we take this image or a few images that show this pattern, and we have a mechanism, a machine learning mechanism, by which we can ask the fleet to source us examples that look like that. The fleet might respond with images that contain those patterns. As an example, these six images might come from the fleet. They all contain bikes on backs of cars.

We would go in, and we would annotate all those as just a single car, then the performance of that detector actually improves, and the network internally understands that, hey, when the bike is just attached to the car, that's actually just a single car. It can learn that given enough examples.

That's how we sort of fixed that problem. I will mention that I talked quite a bit about sourcing data from the fleet. I just want to make a quick point that we've designed this from the beginning with privacy in mind, and all the data that we use for training is anonymized. The fleet doesn't just respond with bicycles on backs of cars. We look for lots of things all the time. For example, we look for boats, and the fleet can respond with boats.

We look for construction sites, the fleet can send us lots of construction sites from across the world. We look for even slightly more rare cases. For example, finding debris on the road is pretty important to us. These are examples of images that have streamed to us from the fleet that show tires, cones, plastic bags, and things like that.

If we can source these at scale, we can annotate them correctly, and the neural network can learn how to deal with them in the world. Here's another example. Animals, of course, also a very rare occurrence and event, we want the neural network to really understand what's going on here, that these are animals and we want to deal with that correctly. To summarize, the process by which we iterate on neural network predictions looks something like this.

We start with a seed dataset that was potentially sourced at random. We annotate that dataset, we train neural networks on that dataset and put that in the car. We have mechanisms by which we notice inaccuracies in the car when this detector may be misbehaving.

For example, if we detect that the neural network might be uncertain or if there's a driver intervention on any of those settings, we can create this trigger infrastructure that sends us data of those inaccuracies. For example, if we don't perform very well on lane line detection on tunnels, we can notice that there's a problem in tunnels. That image would enter our unit test so we can verify that we're actually fixing the problem over time.

Now what you do is to fix this inaccuracy, you need to source many more examples that look like that. We ask the fleet to please send us many more tunnels, we label all those tunnels correctly. We incorporate that into the training set, we retrain the network, redeploy, and iterate the cycle over and over again.

We refer to this iterative process by which we improve these predictions as the data engine. Iteratively deploying something, potentially in shadow mode, sourcing inaccuracies, and incorporating them into the training set over and over again. We do this basically for all the predictions of these neural networks. So far, I've talked about a lot of explicit labeling. Like I mentioned, we ask people to annotate data.

This is an expensive process in time and also, with respect to. It's just an expensive process. These annotations, of course, can be very expensive to achieve. What I want to talk about also is really to utilize the power of the fleet. You don't want to go through this human annotation bottleneck.

You want to just stream in data and automate it automatically. We have multiple mechanisms by which we can do this. As one example of a project that we recently worked on is the detection of cut-ins. You're driving down the highway, someone is on the left or on the right, and they cut in in front of you into your lane. Here's a video showing the Autopilot detecting that this car is intruding into our lane.

Of course, we'd like to detect a cut-in as fast as possible. The way we approach this problem is we don't write explicit code for is the left blinker on, is the right blinker on, track the car body over time and see if it's moving horizontally. We actually use a fleet learning approach.

The way this works is, we ask the fleet to please send us data whenever they see a car transition from a right lane to the center lane or from left to center. Then what we do is we rewind time backwards, and we automatically can annotate that, hey, that car will in 1.3 seconds cut in in front of you. Then we can use that for training the neural net. The neural net will automatically pick up on a lot of these patterns.

For example, the cars are typically yawed. They're moving this way. Maybe the blinker is on. All that stuff happens internally inside the neural net just from these examples. We ask the fleet to automatically send us all this data. We can get half a million or so images, and all of these would be annotated for cut-ins, then we train the network.

Then we took this cut-in network, and we deployed it to the fleet, but we don't turn it on yet. We run it in shadow mode. In shadow mode, the network is always making predictions. Hey, I think this vehicle is going to cut in from the way it looks. This vehicle is going to cut in. Then we look for mispredictions.

As an example, this is a clip that we had from shadow mode of the cut-in network, It's kind of hard to see, but the network thought that the vehicle right ahead of us on the right was going to cut in. You can sort of see that it's slightly flirting with the lane line. It's sort of encroaching a little bit.

The network got excited, and it thought that that was going to be cut in. That vehicle will actually end up in our center lane. That turns out to be incorrect, and the vehicle did not actually do that. What we do now is we just churn the data engine. We source that ran in the shadow mode. It's making predictions. It makes some false positives, and there are some false negative detections.

We got overexcited sometimes, and sometimes we missed a cut-in when it actually happened. All those create a trigger that streams to us and that gets incorporated now for free. There's no humans harmed in the process of labeling this data, incorporated for free into our training set. We retrained the network and redeployed the shadow mode.

We can spin this a few times, and we always look at the false positives and negatives coming from the fleet. Once we're happy with the false positive, false negative ratio, we actually flip the bit and actually let the car control to that network. You may have noticed we actually shipped one of our first versions of a cut-in detector, approximately, I think, three months ago. If you've noticed that the car is much better at detecting cut-ins, that's fleet learning operating at scale.

Yes, it actually works quite nicely. That's fleet learning. No humans were harmed in the process. It's just a lot of neural network training based on data and a lot of shadow mode and looking at those results.

Elon Musk
CEO, Tesla

Essentially like, everyone's training the network all the time is what it amounts to. Whether the Autopilot is on or off, the network is being trained. Every mile that's driven, for the car that's 102 or above, is training the network.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah. Another interesting way that we use this in the scheme of fleet learning, and the other project that I will talk about is a path prediction. While you are driving the car, what you're actually doing is you are annotating the data because you are steering the wheel. You're telling us how to traverse different environments.

What we're looking at here is some person in the fleet who took a left through an intersection. What we do here is we have the full video of all the cameras, and we know that the path that this person took, because of the GPS, the inertial measurement unit, the wheel angle, the wheel ticks. We put all that together, and we understand the path that this person took through this environment. Of course, we can use this for supervision for the network.

We just source a lot of this from the fleet. We train a neural network on those trajectories, the neural network predicts paths just from that data. Really what this is referred to typically is called imitation learning. We're taking human trajectories from the real world, and we're just trying to imitate how people drive in real worlds.

We can also apply the same data engine crank to all of this and make this work over time. Here's an example of path prediction going through a kind of a complicated environment. What you're seeing here is a video, and we are overlaying the predictions of the network. This is a path that the network would follow, in green.

Elon Musk
CEO, Tesla

I mean, the crazy thing is the network is predicting paths it can't even see with incredibly high accuracy. It can't see around the corner, it's saying the probability of that curve is extremely high, that's the path. It nails it. You will see that in the cars today. We're going to turn on augmented vision so you can see the lane lines and the path predictions of the cars overlaid on the video.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah. There's actually more going on under the hood that you can even tell.

Elon Musk
CEO, Tesla

It's kind of scary, to be honest. Yeah.

Andrej Karpathy
Senior Director of AI, Tesla

Of course, there's a lot of details I'm skipping over. You might not want to annotate all the drivers. You might want to just imitate the better drivers, and there's many technical ways that we actually slice and dice that data. The interesting thing here is that this prediction is actually a 3D prediction that we project back to the image here.

The path here forward is a three-dimensional thing that we're just rendering in 2D. But we know about the slope of the ground from all this, and that's actually extremely valuable for driving. Path prediction actually is live in the fleet today, by the way. If you're driving cloverleafs, if you're in a cloverleaf on a highway, until maybe five months ago or so, your car would not be able to do cloverleaf. Now it can. That's path prediction running live on your cars. We've shipped this a while ago.

Today you are going to get to experience this for traversing intersections. A large component of how we go through intersections in your drives today is all sourced from path prediction from automatic labels. What I talked about so far is really the three key components of how we iterate on the predictions of the network and how we make it work over time. You require large, varied, and real dataset.

We can really achieve that here at Tesla and we do that through the scale of the fleet, the data engine, shipping things in shadow mode, iterating that cycle, and potentially even using fleet learning where no human annotators are harmed in the process, and just using data automatically, and we can really do that at scale.

In the next section of my talk, I'm going to especially talk about depth perception using vision only. You might be familiar that there are at least two sensors in a car. One is vision, cameras just getting pixels, and the other is LiDAR, that a lot of companies also use. LiDAR gives you these point measurements of distance around you. One thing I'd like to point out first of all is you all came here-- you drove here, many of you, and you used your neural net and vision. You were not shooting lasers out of your eyes, and you still ended up here.

Elon Musk
CEO, Tesla

We might have.

Andrej Karpathy
Senior Director of AI, Tesla

Things went well.

Elon Musk
CEO, Tesla

Has to be for everyone.

Andrej Karpathy
Senior Director of AI, Tesla

Clearly the human neural net derives distance and all the measurements and the 3D understanding of the world just from vision. It actually uses multiple cues to do so. I'll just briefly go over some of them just to give you a sense of roughly what's going on inside. As an example, we have two eyes pointed out.

You get two independent measurements at every single time step of the world ahead of you. Your brain stitches this information together to arrive at some depth estimation because you can triangulate any points across those two viewpoints. A lot of animals instead have eyes that are positioned on the sides, so they have very little overlap in their visual fields.

They will typically use structure for motion, the idea is that they bob their heads, because of the movement, they actually get multiple observations of the world, you can triangulate again depth. Even with one eye closed and completely motionless, you can still have some sense of depth perception.

If you did this, I don't think you would notice me coming two meters towards you or 100 m back. That's because there are a lot of very strong monocular cues that your brain also takes into account. This is an example of a pretty common visual illusion where you have these two blue bars are identical, but your brain, the way it stitches up this scene, is it just expects one of them to be larger than the other because of the vanishing lines of this image.

Your brain does a lot of this automatically. Neural nets, artificial neural nets can as well. Let me give you three examples of how you can arrive at depth perception from vision alone, a classical approach and two that rely on neural networks. Here's a video going down, I think this is San Francisco, of a Tesla.

These are our cameras, our sensing, and we're looking at all-- I'm only showing the main camera, but all the cameras are turned on, the eight cameras that they are about. If you just have this six-second clip, what you can do is you can stitch up this environment in 3D using multi-view stereo techniques. This Oops. This is supposed to be a video. Is it not a video? Oh, I know it's-

Elon Musk
CEO, Tesla

Oh, there we go.

Andrej Karpathy
Senior Director of AI, Tesla

There we go. This is the 3D reconstruction of those six seconds of that car driving through that path. You can see that this information is very well recoverable from just videos. Roughly that's through process of triangulation and as I mentioned, multi-view stereo, and we've applied similar techniques on a slightly more sparse and approximate also in the car.

It's remarkable all that information is really there in the sensor and just a matter of extracting it. The other project that I want to briefly talk about is, as I mentioned, Neural networks are very powerful visual recognition engines, and if you want them to predict depth, then you need to, for example, look for labels of depth, and then they can actually do that extremely well.

There's nothing limiting networks from predicting this monocular depth except for labeled data. One example project that we've actually looked at internally is we use the forward-facing radar, which is shown in blue, and that radar is looking out and measuring depths of objects. We use that radar to annotate what vision is seeing, the bounding boxes that come out of the neural networks.

Instead of human annotators telling you, okay, this car and this bounding box is roughly 25 m away, you can annotate that data much better using sensors. You use sensor annotation. As an example, radar is quite good at that distance. You can annotate that, and then you can train a neural network on it. If you just have enough data of it, this neural network is very good at predicting those patterns.

Here's an example of predictions of that. In circles, I'm showing radar objects, and the cuboids that are coming out here are purely from vision. The cuboids here are just coming out of vision, and the depth of those cuboids is learned by a sensor annotation from the radar.

If this is working very well, then you would see that the circles in the top-down view would agree with the cuboids, and they do. That's because neural networks are very competent at predicting depths. They can learn the different sizes of vehicles internally, and they know how big those vehicles are, and you can actually derive depth from that quite accurately.

The last mechanism I will talk about very briefly is slightly more fancy and gets a bit more technical, but it is a mechanism that has recently. There's a few papers basically over the last year or two on this approach. It's called self-supervision. What you do in a lot of these papers is you only feed raw videos into neural networks with no labels whatsoever, and you can still get neural networks to learn depth.

It's a little bit technical, so I can't go into the full details, but the idea is that the neural network predicts depth at every single frame of that video, and then there are no explicit targets that the neural network is supposed to regress to with the labels. Instead, the objective for the network is to be consistent over time. Whatever depths you predict should be consistent over the duration of that video. The only way to be consistent is to be right.

The neural network automatically predicts the correct depths for all the pixels. We've reproduced some of these results internally, this also works quite well. In summary, people drive with vision only. No lasers are involved. This seems to work quite well. The point that I'd like to make is that visual recognition and very powerful visual recognition is absolutely necessary for autonomy. It's not a nice-to-have.

Like we must have neural networks that actually really understand the environment around you. LiDAR points are much less information-rich environment. Vision really understands the full details. Just a few points around, there's much less information in those. As an example, on the left here, is that a plastic bag or is that a tire?

A LiDAR might just give you a few points on that, but vision can tell you which one of those two is true. That impacts your control. Is that person who is slightly looking backwards, are they trying to merge into your lane, on the bike, or are they just going forward? In the construction sites, what do those signs say? How should I behave in this world?

The entire infrastructure that we have built up for roads is all designed for human visual consumption. All the signs, all the traffic lights, everything is designed for vision, and so that's where all that information is. You need that ability. Is that person distracted and on their phone? Are they going to walk into your lane? Those answers to all these questions are only found in vision and are necessary for level 4, level 5 autonomy.

That is the capability that we are developing at Tesla. This is done through combination of large scale neural network training, through data engine and getting that to work over time and using the power of the fleet. In this sense, LiDAR is really a shortcut. It sidesteps the fundamental problems, the important problem of visual recognition that is necessary for autonomy. It gives a false sense of progress and it's ultimately a crutch. It does give really fast demos.

If I was to summarize my entire talk in one slide, it would be this. All of autonomy, because you want level 4, level 5 systems that can handle all the possible situations in 99.99% of the cases. Chasing some of the last few nines is going to be very tricky and very difficult and is going to require a very powerful visual system.

I'm showing you some images of what you might encounter in any one slice of that nine. In the beginning, you just have very simple cars going forward. Those cars start to look a little bit funny. Maybe you have bikes and cars. Maybe you have cars and cars. Maybe you start to get into really rare events like cars turned over or even cars airborne. We see a lot of things coming from the fleet, and we see them at some rate, at a really good rate compared to all of our competitors.

The rate of progress at which you can actually address these problems, iterate on the software, and really feed the neural networks with the right data, that rate of progress is really just proportional to how often you encounter these situations in the wild, and we encounter them significantly more frequently than anyone else, which is why we're going to do extremely well. Thank you.

Elon Musk
CEO, Tesla

Thanks. Go ahead.

Colin Rusch
Analyst, Oppenheimer

It's all super impressive. Thank you so much. How many pictures are you collecting on average from each car per period of time? It sounds like the new hardware, with the dual-dual, active-active computers, gives you some really interesting opportunities to run in full simulation, one copy of the neural net while you're running the other one, drive the car and compare the results to do quality assurance.

I was also wondering if there are other opportunities to use the computers for training, when they're parked in the garage for the 90% of the time that I'm not driving my Tesla around. Thank you very much.

Andrej Karpathy
Senior Director of AI, Tesla

Yep. For the first question, how much data do we get from the fleet? It's really important to point out, it's not just the scale of the dataset, it really is the variety of that dataset that matters. If you just have lots of images of something going forward on the highway a t some point, the neural net just gets it. You don't need that data.

Colin Rusch
Analyst, Oppenheimer

Yeah.

Andrej Karpathy
Senior Director of AI, Tesla

We are really strategic in how we pick and choose, and the trigger infrastructure that we've built up is quite sophisticated and allows us to get just the data that we need right now. It's not a massive amount of data, it's just very well-picked data. For the second question, with respect to redundancy, absolutely. You can run basically the copy of the network on both, and that is actually how it's designed, to achieve a level 4, level 5 system that is redundant. That's absolutely the case. Your last question, I'm sorry, I did not.

Elon Musk
CEO, Tesla

Training.

Andrej Karpathy
Senior Director of AI, Tesla

Training.

Elon Musk
CEO, Tesla

The car is an inference-optimized computer. We do have a major program at Tesla, which we don't have enough time to talk about today, called Dojo. That's a super powerful training computer. The goal of Dojo will be to be able to take in vast amounts of data and train at a video level, and do unsupervised massive training of vast amounts of video with the Dojo program, Dojo computer. That's for another day.

Colin Rusch
Analyst, Oppenheimer

All right. I'm an avid user of Autopilot on a daily basis, I guess you could say I'm like a test pilot in a way because I drive the 405, 10 and all these really tricky, really long tail things happen every day. The one challenge that I'm curious to how you're going to solve is changing lanes because whenever I try to get into a lane with traffic, everybody cuts you off. Human behavior is very irrational when you're driving in L.A., the car just wants to do it safely. You almost have to do it unsafely. I was wondering how you're going to solve that problem.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah. One thing I will point out is I spoke about the data engine as iterating on neural networks, but we do the exact same thing on the level of software and all the hyperparameters that go into the choices of when we actually lane change, how aggressive we are. We're always changing those, potentially running them in shadow mode and seeing how well they work.

To tune our heuristics around when it's okay to lane change, we would also potentially utilize the data engine and the shadow mode and so on. Ultimately, actually designing all the different heuristics for when it's okay to lane change is actually a little bit intractable, I think, in the general case. Ideally, you actually want to use fleet learning to guide those decisions. When do humans lane change?

In what scenarios, and when do they feel it's not safe to lane change? Let's just look at a lot of the data and train machine learning classifiers for distinguishing when it is safe to do so. Those machine learning classifiers can write much better code than humans because they have the massive amount of data backing that. They can really tune all the right thresholds and agree with humans and do something safe.

Elon Musk
CEO, Tesla

Well, I think we'll probably have a mode that goes beyond Mad Max mode to L.A. traffic mode. I'm a Mad Max mode. Yeah. Well, Mad Max would have a hard time in L.A. traffic, I think.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah. It really is a trade-off. You don't want to create unsafe situations, you want to be assertive. That little dance of how you make that work as a human is actually very complicated, and it's very hard to write in code. I think it really does seem like machine learning approach is the right way to go about it, where we just look at a lot of ways that people do this and try to imitate that.

Elon Musk
CEO, Tesla

We're just being more conservative right now, as we gain higher and higher confidence, we'll allow users to select a more aggressive mode. That'll be up to the user. In the more aggressive modes, in trying to merge in traffic, no matter how many there's a slight chance of a fender bender. Not a serious accident.

You basically will have a choice of do you want to have a non-zero chance of a fender bender on freeway traffic, which is unfortunately the only way to navigate L.A. traffic. Yes. There's also that. Yeah. You do have to take some risks and crashes between lanes. Yes. You know? Yes. It always reminds me of L.A. Story, which is a great movie.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah, it's very solid because this is a game of chicken that's going on.

Elon Musk
CEO, Tesla

I like that we're talking about driving. Yeah. At some point, I got to say no to slight chance. We'll offer more aggressive options over time that will be user-specified. Mad Max Plus. Yes, Mad Max Plus, exactly.

Andrej Karpathy
Senior Director of AI, Tesla

Oh, yeah. Please.

Jed Dorsheimer
Analyst, Canaccord Genuity

Hello. Hi, Jed Dorsheimer from Canaccord Genuity. Thank you and congratulations on everything that you've developed. When we look at the AlphaZero project, it was a very defined and limited variable in terms of the parameters on that, which allowed for the learning curve to be so quick. What you're trying to do here is almost develop consciousness in the cars through the neural network.

I guess the challenge is how do you not create a circular reference in terms of the pulling from the centralized model of the fleet, to that handoff where the car has enough information? Where is that line, I guess, in terms of the point of the learning process to handing it off, where there's enough information in the car and not having to pull from the fleet?

Elon Musk
CEO, Tesla

Well, the car can operate if it's completely disconnected from the fleet. It uploads the training that's better and better as the fleet gets better and better. Simply, if you disconnected it from the fleet, from that point onwards, it would stop getting better. It would still function fine.

Jed Dorsheimer
Analyst, Canaccord Genuity

I guess in the hardware portion of your share, in the previous version, you talked about a lot of the power benefits of not storing a lot of the images. In this portion, you're talking about the learning that's going on by pulling from the fleet.

I guess I'm having a hard time reconciling how if there was a situation where I'm driving up the hill as you showed, and I'm predicting where the road is going to go, that's coming from all of the other fleet variables that led to that intelligence. How I'm getting the benefit of the low power using the cameras with the neural network. That's where I'm losing the two. Maybe it's just me, I guess that's-

Elon Musk
CEO, Tesla

The compute power in the Full Self-Driving computer is incredible. Maybe we should mention that if it had never seen that road before, it would still have made those predictions, provided it was a road in the U.S.

Adam Jonas
Analyst, Morgan Stanley

I n the case of LiDAR, the March of Nines, isn't there an example? I want to just get to your slam on LiDAR because it's pretty clear you don't like LiDAR.

Elon Musk
CEO, Tesla

LiDAR is lame. LiDAR is lame. Damn.

Adam Jonas
Analyst, Morgan Stanley

Isn't there a case where at some point, 99999 down the road, where actually LiDAR may be helpful, and why not have it as some sort of a redundancy or backup? That's my first question. The second, you can still have your focus on computer vision, but just have it as a redundant. My second question is: If that is true, what happens to the rest of the industry that's building their autonomy solutions on LiDAR?

Elon Musk
CEO, Tesla

They're all going to dump LiDAR, is my prediction. Mark my words. I should point out that I don't actually super hate LiDAR as much as it may sound. SpaceX Dragon uses LiDAR to navigate to the space station and dock. Not only that, SpaceX developed its own LiDAR from scratch to do that, and I spearheaded that effort personally. In that scenario, LiDAR makes sense.

In cars, it's freaking stupid. It's expensive and unnecessary, and as Andrej was saying, once you solve vision, it's worthless. You have expensive hardware that's worthless on the car. We do have a forward radar, which is low cost and is helpful, especially for occlusion situations. If there's fog or dust or snow, the radar can see through that.

If you're going to use active photon generation, don't use visible wavelength, because with passive optical, you've taken care of all visible wavelength stuff. You want to use a wavelength that is occlusion-penetrating like radar. LiDAR is just active photon generation in the visual spectrum.

If you're going to do active photon generation, do it outside the visual spectrum in the radar spectrum. Like at 3.8 mm versus 400 nm - 700 nm, you're going to be much better occlusion penetration, and that's why we have a forward radar. We also have 12 ultrasonics for near-field information, in addition to the eight cameras and the forward radar. You only need the radar in the forward direction because that's the only direction you're going real fast. I mean, we've gone over this multiple times, like are we sure we have the right sensor suite? Should we add anything more? No.

Speaker 16

Hey, Elon. Elon.

Elon Musk
CEO, Tesla

Okay.

Speaker 16

Hi. Right here. You had mentioned that you asked the fleet for the information that you're looking for some of the vision. I have two questions about that. It sounds like the cars are doing some computation to determine what kind of information to send back to you. Is that a correct assumption?

Elon Musk
CEO, Tesla

Yeah.

Speaker 16

Are they doing that in real time, or are they doing based on stored information?

Andrej Karpathy
Senior Director of AI, Tesla

Yep. They absolutely do computation in real time on the car. We have a way to basically specify condition that we're interested in, and then those cars do that computation there. If they did not, then we'd have to send all the data and do that offline in our backend. We don't want to do that. All that computation happens on the car.

Speaker 16

Based on that question, it sounds like you guys are in a really good position to have currently half a million cars, in the future, potentially millions of cars t hat are essentially computers representing almost free data centers for you-

Elon Musk
CEO, Tesla

Yes.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah

Speaker 16

... to do computational. Is that a huge future opportunity for Tesla?

Elon Musk
CEO, Tesla

It's current.

Speaker 16

A current opportunity. That's not really factored in for anything yet. That's incredible. Thank you.

Elon Musk
CEO, Tesla

We have 425,000 cars with Hardware 2 and beyond, which means they've got all eight cameras, the radar, and ultrasonics. They've got at least the NVIDIA computer, which is enough to essentially figure out what information is important, what is not, compress the information that is important to the most salient elements, and upload it to the network for training. It's a massive compression of real-world data.

Speaker 16

Do you see ever using that computational capacity for other things besides self-driving? You have these sort of network of millions of computers, which is like massive data centers essentially that are distributed data centers for computational capacity. Do you see it being used for other things besides self-driving in the future?

Elon Musk
CEO, Tesla

I suppose it could possibly be used for something besides self-driving. We've been super focused on self-driving, so as we get that really nailed, maybe there's going to be some other use for millions and then tens of millions of computers with Hardware 3 or a Full Self-Driving computer. Yeah, maybe there would be.

Speaker 16

Another question on energy storage from them as well.

Elon Musk
CEO, Tesla

It could be. Maybe there's some sort of AWS angle here. It's possible.

Matt Joyce
Analyst, Loup Ventures

Hello. Hi, Elon. Matt Joyce, Loup Ventures. I own a Model 3 in Minnesota, where it snows a lot. Since camera and radar cannot see road markings through snow, what is your technical strategy to solve this challenge? Does it involve high-precision GPS at all?

Elon Musk
CEO, Tesla

Go ahead.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah. Actually, like today, actually, Autopilot will do a decent job in snow, even when lane markings are covered. Even when lane markings are faded, covered, or when there's lots of rain on them, we still seem to drive relatively well. We didn't specifically go after snow yet with our data engine. I actually think this is completely tractable because in a lot of those images, even when things are snowy, when you ask a human annotator, where are the lane lines? They actually could tell you. They actually are-

Elon Musk
CEO, Tesla

Most likely.

Andrej Karpathy
Senior Director of AI, Tesla

... relatively consistent-

Elon Musk
CEO, Tesla

Yeah

Andrej Karpathy
Senior Director of AI, Tesla

... in creating those lane lines. As long as the annotators are consistent on your data, then the neural network will pick up on those patterns, and we'll do just fine. It's really just about, is the signal there even for the human annotator? If the answer to that is yes, then the neural network can do it just fine.

Elon Musk
CEO, Tesla

Yeah. There are a number of important signals, as Andrej was saying. Lane lines are one of those things, but one of the most important signals is drive space. What is drivable space and what is not drivable space? What actually really matters the most is drivable space more than lane lines. The prediction of drivable space is extremely good. I think especially after this upcoming winter will be incredible. It will be like, how could it possibly be that good? That's crazy.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah.

Elon Musk
CEO, Tesla

Yeah.

Andrej Karpathy
Senior Director of AI, Tesla

The other thing to point out is, maybe it's not even only about human annotators. As long as you as a human can drive through that environment-

Elon Musk
CEO, Tesla

Yeah, exactly.

Andrej Karpathy
Senior Director of AI, Tesla

... through fleet learning, we actually know the path you took. You obviously use vision to guide you through that path. You did not just use the lane line markings. You used the entire geometry of the entire scene. You see how the road is roughly curling. You see how the cars are positioned around you. The neural network will pick up on all those patterns automatically and see it if you just have enough of the data of people traversing those environments.

Elon Musk
CEO, Tesla

Yeah. It's actually extremely important that things not be rigidly tied to GPS because GPS error can vary quite a bit. The actual situation for a road can vary quite a bit. There could be construction, there could be a detour. If a car is using GPS as primary, this is a real bad situation. It's asking for trouble. It's fine to use GPS for tips and tricks.

It's like you can drive your home neighborhood better than a neighborhood in some other country or some other part of the country. You know your own neighborhood well, and you use the knowledge of your own neighborhood to drive with more confidence, to maybe have counterintuitive shortcuts and that kind of thing. The GPS overlay data should only be helpful, but never primary. If it's ever primary, it's a problem.

Speaker 16

Question back here in the back corner. I just wanted to follow up partially on that, because several of your competitors in the space over the past few years have talked about how they are augmenting all of their perception and path planning capabilities that are on the car platform with high-definition maps of the areas that they are driving. Does that play a role in your system? Do you see it adding any value? Are there areas where you would like to get more data that is not collected from the fleet but is more kind of mapping style data?

Elon Musk
CEO, Tesla

I think the high-precision GPS maps and lanes are a really bad idea. The system becomes extremely brittle. Any change to the system makes it can't adapt. If it locks onto GPS and high-precision lane lines, and does not allow vision override. In fact, vision should be the thing that does everything. Then lane lines are a guideline, but they're not the main thing. We briefly balked up the tree of high-precision lane lines, realized that was a huge mistake and reversed it out. It's not good.

Speaker 16

This is very helpful for understanding annotation, where the objects are, and how the car drives. What about the negotiation aspect for parking and roundabouts and other things where there are other cars on the road that are human-driven, where it's more art than science?

Elon Musk
CEO, Tesla

It does pretty good, actually. Like with cut-ins and stuff, it's doing really well.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah. Like I mentioned, we're using a lot of machine learning right now in terms of predicting, kind of creating an explicit representation of what the role looks like. Then there's an explicit planner and a controller on top of that representation. There's a lot of heuristics for how to traverse and negotiate and so on.

There is a long tail, just like in what the visual environments look like. There's a long tail in just those negotiations and the little game of chicken that you play with other people and so on. I think we have a lot of confidence that eventually there must be some kind of a fleet learning component to how you actually do that, because writing all those rules by hand is going to quickly plateau, I think.

Elon Musk
CEO, Tesla

Yeah. We've dealt with this issue with cut-ins, and it's like we'll allow gradually more aggressive behavior on the part of the user. They can just dial the setting up and say, be more aggressive, be less aggressive, drive easy, chill mode, aggressive. Yeah.

Trip Chowdhry
Analyst, Global Equities Research

Incredible progress. Phenomenal. Two questions. First, in terms of platooning, do you think the system is geared? Somebody asked about when there is snow on the road, but if you have platooning feature, you can just follow the car in front. Is your system capable of doing that? I have two follow-ups.

Andrej Karpathy
Senior Director of AI, Tesla

You're asking about platooning. I think we could absolutely build those features. Again, if you just train neural networks, for example, on imitating humans already follow the car ahead. That neural network actually incorporates those patterns internally. It figures out that there's a correlation between the way the car ahead of you faces and the path that you are going to take, but that's all done internally in the net.

You're just concerned with getting enough data and the tricky data, and the neural network training process actually is quite magical, does all the other stuff automatically. You turn all the different problems into just one problem. Just collect your dataset and use neural network training.

Elon Musk
CEO, Tesla

There's three steps to self-driving. There's being feature complete, then there's being feature complete to the degree where we think that the person in the car does not need to pay attention, and then there's being at a reliability level where we've also convinced regulators that that is true. There's kind of three levels.

We expect to be feature complete in self-driving this year. We expect to be confident enough from our standpoint to say that we think people do not need to touch the wheel or look out of the window sometime probably around, I don't know, second quarter of next year. We expect to get regulatory approval, at least in some jurisdictions, for that towards the end of next year. That's roughly the timeline that I expect things to go on.

Probably for trucks, the platooning will be approved by regulators before anything else. You could have, maybe if you're doing long-haul freight, you could have one driver in the front and then have four semis trailing behind in a platooning manner. I think that probably regulators will be quicker to approve that than other things.

Trip Chowdhry
Analyst, Global Equities Research

Regarding, of course, you don't have to convince us that LiDAR is a technology, in my opinion, which has an answer looking for a question probably dead. This is very impressive what we saw today, probably a demo could show something more. I was wondering, what is the maximum dimension of a matrix that you may be having in your training or in your deep learning pipeline? A ballpark figure.

Andrej Karpathy
Senior Director of AI, Tesla

Maximum dimension of the matrix. Doing a lot of matrix multiply operations inside the neural network, you're asking about There's many different ways to answer that question, but I'm not 100% sure if they're useful answers. These neural networks will typically have, like I mentioned, about tens to hundreds of millions of neurons. Each of them on average have about 1,000 connections to neurons below. Those are the typical scales that are used across the industry, and also that we would use as well.

Speaker 16

Yeah. I've been actually very impressed by the rate of improvement on Autopilot the past year on my Model 3. The two scenarios I wanted your feedback on. Last week, the first scenario was, I was on the right-hand most lane of the freeway, and there was a highway on-ramp, and then my Model 3 actually was able to detect the two cars on the side, slow down, and let the car go in front of me and one car go behind me, and I was like, oh, my gosh. This is insane. I didn't think my Model 3 could do that. That was super impressive.

The same week, another scenario, which is I was on the right-hand lane again, but my right-hand lane was merging with the left lane. It wasn't an on-ramp. It's just a normal highway freeway lane. My Model 3 wasn't able to detect really that situation, and I wasn't able to slow down or speed up, and I had to intervene. Can you, from your perspective, share the background on how a neural net how Tesla might adjust for that? How that could be improved over time.

Andrej Karpathy
Senior Director of AI, Tesla

Yeah. Like I mentioned, we have a very sophisticated trigger infrastructure. If you have intervened, it's actually potentially likely that we receive that clip and that we can actually analyze it and see what happened and tune the system. It probably enters some statistics over, at what rate are we correctly merging the traffic? We look at those numbers, and we look at the clips, and we see what's wrong, and we try to fix those clips and make progress against those benchmarks. Yeah.

Speaker 16

Is it like a manual vision thing, like you have people looking at video?

Andrej Karpathy
Senior Director of AI, Tesla

Yeah. We would potentially go through a phase of categorization, and then we look at some of the biggest categories that actually seem to semantically be related to the same problem, and then we would look at some of those and then try to develop software against that.

Elon Musk
CEO, Tesla

Okay. We do have one more presentation, which is the software. There's essentially the Autopilot hardware, with Stuart, there's the neural net vision, with Andrej, and then there's the software engineering at scale, that's going to be presented by Stuart. Thanks. T here'll be an opportunity afterwards to ask questions. Yeah, thanks.

Martin Viecha
Director of Investor Relations, Tesla

I just wanted to very briefly say, if you have an early flight and you want to do a test ride with our latest development software, if you could please speak to my colleague, Anne, or drop her an email, and we can take you out for a test ride. Stuart, over to you.

Stuart Bowers
VP of Engineering, Tesla

That's actually from a clip of a longer than 30-minute uninterrupted drive, with no interventions, Navigate on Autopilot on the highway system, which is in production today on hundreds of thousands of cars. I'm Stuart, and I'm here to talk about how we build some of these systems at scale. Just like a really short introduction on where I'm coming from-

Elon Musk
CEO, Tesla

Yeah

Stuart Bowers
VP of Engineering, Tesla

... what I do. I've been in a couple of companies more or less. I've been writing software professionally for about 12 years. The thing that excites me most and I'm really passionate about is taking the cutting edge of machine learning and actually connecting that with customers through robustness and scale. At Facebook, I worked initially inside of our ads infrastructure to build some of the machine learning. Some really, really smart people.

We actually tried to build that into a single platform that we could then scale to all the other aspects of the business, from how we rank the news feeds, to how we deliver search results, to how we make every recommendation across the platform. That became the applied machine learning group, and that's something which I'm incredibly proud of.

A lot of that wasn't just the core algorithms and the really important improvements that happened there, though that matters. A lot of it is actually the engineering practices to build these systems at scale. The same thing was true at Snap where I went, where we were really, really excited to sort of actually help to monetize this product.

The hardest part, we were using Google at the time, and they were effectively running us on a very small scale, and we wanted to build that same infrastructure. We take an understanding of these users, connect that with cutting-edge machine learning, build that at massive scale, and handle billions and then trillions of both predictions and auctions every day, in a way which is really robust.

When the opportunity came to come to Tesla, that's something I'm just like incredibly excited to do, which is specifically take the amazing things that are happening both in the hardware side and the computer vision and AI side, and actually package that together with all the planning, the controls, the testing, the kernel patching of the operating system, all of our continuous integration, our simulation, and actually build that into a product we get onto people's cars in production today.

I want to talk about the timeline for how we did that with Navigate on Autopilot and how we're going to do that as we get Navigate on Autopilot off the highway and onto city streets. We're at 70 million miles already for Navigate on Autopilot. It's something really, really cool. I think one thing that is worth kind of calling out on this is that we're continuing to accelerate and keep learning from this data. Andrej talked about this data engine. This accelerates up, we actually do make more a nd more assertive lane changes.

We are learning from these cases where people intervene, either because they failed to detect a merge correctly or because they wanted the car to be a little more peppy in different environments, we just want to keep making that progress. To start all of this, we begin with trying to understand the world around us. We talked about the different sensors in the vehicle, but I want to dig in a little bit more here. We have eight cameras, but then we also have additionally 12 ultrasonic sensors, a radar, an inertial measurement unit, GPS.

One thing we forget about is we also have the pedal and steering actions. Not only can we look at what's happening around the vehicle, we can look at how humans chose to interact with that environment. I'll talk to this clip right now. This basically is showing what's happening today in the car, we're continuing to push this forward. We start with a single neural network. We see the detections around it.

We build all that together with multiple neural networks and multiple detections. We bring in the other sensors, we convert that into what Elon calls a vector space, an understanding of the world around us. This is something where as we continue to get better and better at this, we're moving more and more of this logic into the neural networks themselves.

The obvious end game here is that the neural network looks across all the cars, brings in all the information together, and just ultimately outputs a source of truth for the world around us. This is actually not like an artist rendering in many senses. This is actually the output of one of the debugging tools that we use on the team every day to understand what the world looks like around us.

Another thing that I think is really, really exciting to me, and I think when I do hear about sensors like LiDAR, a common question is around just having extra sensor modalities. Why not have some redundancy on the vehicle? I want to dig in on one thing that is not always obvious with neural networks themselves. We have a neural network running on our, say, wide fisheye camera.

That neural network is not making one prediction about the world. It's making many separate predictions, some of which actually audit each other. As a real example, we have the ability to detect a pedestrian. That's something we train very, very carefully on and put a lot of work into. We also have the ability to detect obstacles in the roadway, and a pedestrian is an obstacle, and it's shown differently to the neural network.

It says, oh, there's a thing I can't drive through. These together combine to give us an increased sense of what we can and can't do in front of the vehicle and how to plan for that. We do this across multiple cameras because we have overlapping fields of view in many places around the vehicle. In front, we have a particularly large number of overlapping fields of view.

Lastly, we can combine that with things like the radar and the ultrasonics to build these extremely precise understandings of what's happening in front of the car. We can use that both to learn future behaviors that are very accurate, but we can also build very accurate predictions of how things will continue to happen in front of us.

One example I think is really exciting is we can actually look at bicyclists and people and not just ask where are you now, but where are you going? This is actually the heart of what we're doing for our next generation automatic emergency braking system, which will not just stop for people in your path, but it'll stop for people who are going to be in your path.

That's running in shadow mode right now, will go out to the fleet this quarter. I'll talk about shadow mode in a second. When you want to start a feature like this for Navigate on Autopilot on the highway system, you can start by learning from data. You can just look at how humans do things today. What is their assertiveness profile?

How do they change lanes? What causes them to either abort or change their maneuvers? You can see things that are not immediately obvious, like, oh, yeah, simultaneous merging is rare, but very complicated and very important. You can start to build opinions about different scenarios, such as a fast overtaking vehicle. This is what we do when we initially have some algorithms we want to try out.

We can put them on the fleet, we can see what they would have done in a real-world scenario, such as this car that's overtaking us very quickly. This is taken from our actual simulation environment, showing different paths that we had considered taking and how those overlay on the real-world behavior of a user.

When you get those algorithms tuned up and you feel good about them specifically, this is really taking that output of the neural network, putting it in that vector space, and building and tuning these parameters on top of it, ultimately, a thing that we can do through more and more machine learning, you go into a controlled deployment, which for us is our Early Access Program.

This is, you get this out to a couple thousand people who are really excited to give you highly vigilant but useful feedback about how this behaves, not in an open loop, but in a closed-loop way in the real world, you watch their interventions. We talked about this, when somebody takes over, we can actually get that clip, try to understand what happens.

One thing we can really do is we can actually play this back, again, in an open-loop way and ask, as we build our software, are we getting closer or further from how humans behave in the real world? One thing which is super cool with the Full Self-Driving computers, we're actually building our own racks and infrastructure.

We basically can fit four of our Full Self-Driving computers fully racked up, build these into our own cluster, actually run this very sophisticated data infrastructure to actually understand over time, as we tune and fix these algorithms, are we getting closer and closer to how humans behave, ultimately, can we exceed their capabilities?

Once we had this and we felt really good about it, we wanted to do our wide rollout. To start, we actually asked everybody to confirm the car's behavior via stalk confirm. We started making lots and lots of predictions about how we should be navigating the highway. We asked people to tell us, is this right or is this wrong? This is again, a chance to turn that data engine.

We did spot some really tricky and interesting long tails of in this case, I think a really fun example, there are these very interesting cases of simultaneous merging, where you start going then somebody moves either behind or before you, not noticing you. What is the appropriate behavior here?

What are the tunings of the neural network we need to do to be super precise about the appropriate behaviors here? We worked, we tuned these in the background. We made them better. Over the course of time, we got 9 million successfully accepted lane changes. We use these, again, with our continuous integration infrastructure, to actually understand, do we think we're ready? This is one thing where Full Self-Driving is also really exciting to me.

Since we own the entire software stack, straight from the kernel patching all the way to the tuning on the image signal processor, we can start to collect even more data that is even more accurate, this allows us to do even better and better tuning at these faster iteration cycles.

Earlier this month, we thought we're ready to deploy an even more seamless version of Navigate on Autopilot on the highway system. That seamless version does not require a stalk confirm. You can sit there, relax, put your hand on the wheel, just oversee what the car is doing. In this case, we're actually seeing over 100,000 automated lane changes every single day on the highway system.

This is something that's just super cool to us to deploy at scale. The thing that I'm most excited about from all this is the actual life cycle of this and how we actually will turn that data engine crank faster and faster and faster with time. I think one thing that's really becoming very clear is the combination of the infrastructure we have built, the tooling we built on top of that, the combined power of the Full Self-Driving computer. I believe we can do this even faster as we move Navigate on Autopilot from the highway system onto city streets. Yeah, with that, I'll hand off to Elon.

Elon Musk
CEO, Tesla

Yeah. To the best of my knowledge, all those lane changes have occurred with zero accidents. Is that correct?

Stuart Bowers
VP of Engineering, Tesla

That is correct, yeah. I watch every single accident.

Elon Musk
CEO, Tesla

Yeah. It's conservative, obviously.

Stuart Bowers
VP of Engineering, Tesla

Sure.

Elon Musk
CEO, Tesla

To have hundreds of thousands going to millions of lane changes and zero accidents is, I think, a great achievement by the Tesla team, yeah.

Stuart Bowers
VP of Engineering, Tesla

Thank you.

Elon Musk
CEO, Tesla

Cool. Let's see. A few other things that are maybe worth mentioning. In order to have a self-driving car or Robotaxi, you really need redundancy throughout the vehicle at the hardware level. Starting in, I believe it was October 2016, all cars made by Tesla have redundant power steering. We have redundant motors in the power steering, if a motor fails, the car can still steer.

All of the power and data lines have redundancy, you can sever any given power line or any data line, and the car will keep driving. The auxiliary power system, even if you lose complete power in the main pack, the car is capable of steering and braking using the auxiliary power system. You can completely lose the main pack, and the car is safe.

The whole system, from a hardware standpoint, has been designed to be a Robotaxi since basically October 2016, when we rolled out Autopilot version two. We do not expect to upgrade cars made before that. We think it would actually cost more to make a new car than to upgrade the cars, just to give you a sense of how hard it is to do this. Unless it's designed in, it's not worth it.

We've gone through the future of self-driving, where it's hardware, it's vision, and then there's a lot of software. The software problem here should not be minimized. It's a massive software problem. Yeah. Managing vast amounts of data, training against the data, how do you control the car based on the vision? It's a very difficult software problem. Going over just Tesla master plan.

Obviously, we've made a bunch of forward-looking statements, as they call it. Let's go through some of our other forward-looking statements that we've made. Way back when we created the company, we said we'd build the Tesla Roadster. They said it was impossible, and that even if we did build it, nobody would buy it. This was universal opinion, was that building an electric car was extremely dumb and would fail.

I agree with them that probably the failure was high, but that this was important. We built the Tesla Roadster, got it into production in 2008, and shipping that car. It's now a collector's item. We said we'd build a more affordable car with the Model S. We did that. Again, we were told that's impossible. I was called a fraud and a liar. It was not going to happen. This was all untrue.

Okay, famous last words. We went into production with the Model S in 2012, exceeded all expectations. There is still, in 2019, no car that can compete with the Model S of 2012. It's seven years later. Still waiting. We pulled an affordable car, maybe highly affordable. It was affordable, more affordable, with the Model 3. We built the Model 3. We're in production. I said we'd get over 5,000 cars a week for Model 3.

At this point, 5,000 cars a week is a walk in the park for us. It's not even hard. Said we'd do large-scale solar, which we did through the SolarCity acquisition, and that we'd develop and deploy the solar roof, which is going really well. We're now on version three of the solar tile roof. We expect to scale up production of the solar tile roof significantly later this year.

I have it on my house, and it's great. I said we'd make the Powerwall and the Powerpack. We made the Powerwall and Powerpack. In fact, the Powerpack is now deployed in massive grid-scale utility systems around the world, including the largest operating battery projects in the world above 100 MW.

In the next, well, probably by next year, two years at the most, we expect to have a gigawatt-scale battery project completed. All these things, I said we would do them, we did it. Said we'd do it, we did it. We're going to do the Robotaxi thing, too. Only criticism, it's a fair one, and sometimes I'm not on time. I get it done, and the Tesla team gets it done.

What we're going to do this year is we're going to reach a combined production of 10,000 a week between S, X, and 3. Feel very confident about that. We feel very confident about being feature complete with self-driving. Next year, we'll expand the product line with Model Y and Semi, and we expect to have the first operating Robotaxis next year with no one in them. Next year.

When things are at an exponential rate of improvement, it's very difficult to wrap one's mind around it because we're used to extrapolating on a linear basis. When you've got massive amounts of hardware on the road, the cumulative data is increasing exponentially. The software is getting better at an exponential rate. I feel very confident predicting autonomous Robotaxis for Tesla next year.

Not in all jurisdictions because we won't have regulatory approval everywhere, but I'm confident we'll have at least regulatory approval somewhere literally next year. Any customer will be able to add or remove their car to the Tesla Network. Expect this to operate as sort of like a combination of maybe the Uber and Airbnb model.

If you own the car, you can add or subtract it to the Tesla Network, and Tesla would take 25% or 30% of the revenue. In places where there aren't enough people sharing their cars, we would just have dedicated Tesla vehicles. When you use the car, we'll show you our ride-sharing app. You're able to summon the car from the parking lot, get in, and go for a drive. It's really simple.

You just take the same Tesla app that you currently have, we'll just update the app and add a Summon Tesla or commit your car to the fleet. See the summon your car or add Summon a Tesla or add or subtract your car to the fleet. You'll be able to do that from your phone. We see potential for smoothing out the demand distribution curve, and having a car operate at a much higher utility than a normal car would operate.

Typically, the use of a car is about 10 to 12 hours a week. Most people will drive one and a half to two hours a day, typically 10 to 12 hours a week of total driving. If you have a car that can operate autonomously, then most likely, you'd have that car operate for a third of the week or longer.

There are 168 hours in a week, probably you've got something on the order of 55 - 60 hours a week of operation, maybe a bit longer. The fundamental utility of a vehicle increases by a factor of five. You look at this from a macroeconomic standpoint and say, if we were operating some big simulation, if you could upgrade your simulation to increase the utility of cars by a factor of five, that would be a massive increase in the economic efficiency of the simulation. Just gigantic.

We'll do Model 3, S, and X as taxis, but we made an important change to our leases. If you lease a Model 3, you don't have the option of buying it at the end of the lease. We want them back. If you buy the car, you can keep it, but if you lease it, you have to give it back. As I said, in any locations where there's not enough supply for sharing, Tesla will just make its own cars and add them to the network in that place.

The current cost of Model 3 Robotaxi is less than $38,000. We expect that number to improve over time. We're designing the cars currently being built are all designed for a million miles of operation. The drive unit's designed and tested and validated for a million miles of operation. The current battery pack is about maybe 300,000 - 500,000 mi. The new battery pack that will probably go into production next year is designed explicitly for a million miles of operation.

The entire vehicle, battery pack inclusive, is designed to operate for a million miles with minimal maintenance. We're actually adjusting tire design and really optimizing the car for a hyper-efficient Robotaxi. At some point, you won't need steering wheels or pedals, and we'll just delete those. As these things become less and less important, we'll just delete parts. They won't be there.

Because say, probably two years from now, we make a car that has no steering wheels or pedals, and if we need to accelerate that time, we can always just delete parts. Easy. Probably, say long term, three years, Robotaxis with eliminated parts, maybe it ends up being $25,000 or less. You want a super efficient car, so the electricity consumption is very low. We're currently at one and a half miles per kilowatt hour, but we'll improve that to five and beyond.

There's just really no company that has the full stack integration. We've got the vehicle design and manufacturing, we've got the computer hardware in-house, we've got the in-house software development, and AI, and we've got by far the biggest fleet. It's extremely difficult, not impossible perhaps, but extremely difficult to catch up when Tesla has 100 x more miles per day than everyone else combined.

These days, this is the cost of running a gasoline car, or the average cost of running a car in the U.S. This is taken from AAA. It's currently about $0.62 a mile. 13,500 mi from 15 million vehicles adds up to $2 trillion a year. These are literally just taken from the AAA website. The cost of ride-sharing is, according to Uber and Lyft, is $2-$3 a mile.

The cost to run a Robotaxi, we think less than $0.18 a mile and dropping. This would be current. This is current cost. Future cost will be lower. If you say, what would be the probable gross profit from a single Robotaxi? We think probably something on the order of $30,000 per year. We're designing the cars the same way that commercial semi-trailers, semi-trucks are designed.

Commercial semi-trucks are all designed for a million-mile life, and we're designing the cars for a million-mile life as well. So in nominal dollars, that would be a little over $300,000 over the course of 11 years, might be higher. I think this consumption is actually relatively conservative, and this assumes that 50% of the miles driven are not useful. This is only at 50% utility.

By the middle of next year, we'll have over a million Tesla cars on the road with Full Self-Driving hardware, feature complete, at a reliability level that we would consider that no one needs to pay attention. Meaning you could go to sleep in your From our standpoint, if you fast-forward a year to maybe a year, maybe a year and three months, but next year for sure, we will have over a million Robotaxis on the road. The fleet wakes up with an over-the-air update. That's all it takes. You say, what is net present value of a Robotaxi? Probably on the order of a couple thousand dollars. Buying a Model 3 is a good deal. Any questions?

Trip Chowdhry
Analyst, Global Equities Research

How many do you want in your fleet? Your own fleet.

Elon Musk
CEO, Tesla

Well, I mean, in our own fleet, I don't know, I guess long-term, we have probably on the order of 10 million vehicles. I mean, our production rates generally, if you look at our compound annual production rate since 2012, that's our first full year of Model S production away from 23,000 vehicles produced in 2013 to around 250,000 produced last year. In the course of five years, we increased output by a factor of 10. I would expect that something similar occurs over the next five or six years. As for sharing versus, I don't know. The nice thing is that essentially customers are fronting us the money for the car. That's great.

Speaker 16

In terms of the, one thing is the snake charger, I'm curious about that. Also, how did you determine the pricing? It looks like you're undercutting the average Lyft or Uber ride by about 50%. I'm curious if you could talk a little bit about the pricing strategy.

Elon Musk
CEO, Tesla

Sure. Solving for the snake charger is pretty straightforward from a vision standpoint. It's like a known situation. Any kind of known situation with vision is, like a charge port, it's trivial. The cars would just automatically park and automatically plug in. There would be no human supervision required. Sorry.

Speaker 16

Oh, pricing. Yeah. We just threw some numbers on there. I think it's definitely plug in whatever pricing you think makes sense. We just kind of randomly said, okay, maybe $1. There's on the order of 2 billion cars and trucks in the world. Robotaxis will be in extremely high demand for a very long time.

From my observation thus far is that the auto industry is very slow to adapt. I mean, like I said, there's still not a car on the road that you can buy today that is as good as the Model S was in 2012. That suggests a pretty slow rate of adaptation for the car industry. Probably $1 is conservative for the next 10 years because people sort of think there's actually not enough appreciation for the difficulty of manufacturing. Manufacturing is insanely difficult.

Elon Musk
CEO, Tesla

A lot of people I talk to think if you just have the right design, you can instantly make as much of that thing as the world wants. This is not true. It's extremely hard to design a new manufacturing system for new technology. Audi's having major problems manufacturing the e-tron, and they are extremely good at manufacturing. If they're having problems, what about others?

There's on the order of 2 billion cars and trucks in the world, on the order of about 100 million units per year of production capacity of vehicles, but only of the old design. It will take a very long time to convert all of that to Full Self-Driving cars. They really need to be electric because the cost of operation of a gasoline diesel car is much higher than an electric car. Any Robotaxi that isn't electric will absolutely not be competitive.

Colin Rusch
Analyst, Oppenheimer

Elon, it's Colin Rusch from Oppenheimer over here.

Elon Musk
CEO, Tesla

Yeah.

Colin Rusch
Analyst, Oppenheimer

Obviously, we appreciate that the customers are fronting some of the cash for this fleet getting built up. It sounds like a massive balance sheet commitment from the organization over the course of time. Can you talk a little bit about what that looks like, what your expectations are in terms of financing over the next, call it three years, three to four years, for building up this fleet and starting to monetize it with your customer base?

Elon Musk
CEO, Tesla

Well, we're aiming to be approximately cash flow neutral during the fleet build-up phase. I would expect to be extremely cash flow positive once the Robotaxis are enabled. I don't want to talk about financing rounds, or be able to talk about financing rounds in this venue. I think we'll make the right moves. I think we'll make the moves we think we should make. Yeah.

Speaker 17

I have a question. If I'm Uber, why wouldn't I just buy all your cars? Why would I let you put me out of business?

Elon Musk
CEO, Tesla

There's a clause that we put into our cars, I think it was about three or four years ago, they can only be used in the Tesla Network.

Speaker 17

Even a private person, like if I go out and buy 10 Model 3s, I can run it on the network, that's a business now, right?

Elon Musk
CEO, Tesla

You're only allowed to use the Tesla Network.

Speaker 17

Right. If I use the Tesla Network, in theory, I could run a car-sharing Robotaxi business with my 10 Model 3s.

Elon Musk
CEO, Tesla

Yes, it's like the App Store. You can only add or remove them, through the Tesla Network, and then Tesla gets a revenue share.

Speaker 17

Similar to Airbnb, though, in that I have this home, my car, and now I can just rent them out. I can make an extra income from owning multiple cars and just renting them out. Like I have a Model 3, I aspire to get this Roadster here next when you build it, and I'm going to just rent my Model 3 out. Why would I give it back to you?

Elon Musk
CEO, Tesla

I guess you could operate a rental car fleet, I think this is very unwieldy. Yeah.

Speaker 17

I don't know. Seems easy.

Elon Musk
CEO, Tesla

Okay. Try it.

Speaker 16

Elon, in order to operate a Robotaxi network, it sounds like you have to solve certain problems. For example, Autopilot today, if you oversteer it lets you take over. If it's a ride-sharing product that someone else is getting in the passenger seat, moving the steering can't let that person take over the car, for example, because they might not even be in the driver's seat.

Is the hardware already there for it to be a Robotaxi, and it might get into situations such as a cop pulling it over where some human might need to intervene, like using central fleet of operators that remotely sort of interact with humans? Is all of that type of infrastructure already built into each of the cars? Does that make sense?

Elon Musk
CEO, Tesla

I think there will be sort of a phone home thing where if the car gets stuck, it'll just phone home to Tesla and ask for a solution. Things like being pulled over by a police officer, that's easy for us to program in. That's not a problem. It will be possible for somebody to take over using the steering wheel, or at least for some period of time, and then probably down the road, we'll just cap the steering wheel so there's no steering control. We'll just take the steering wheel off and put a cap on in the long, give it a couple of years.

Speaker 16

Will that require a hardware modification to the car in order for it to enable that or?

Elon Musk
CEO, Tesla

Yeah, we'll literally just unbolt the steering wheel and put a cap on where the steering wheel handle currently is.

Speaker 16

That is a future car that you would put out. What about today's cars where the steering wheel is a mechanism to take over Autopilot? If it's in a Robotaxi mode, would someone be able to take it over by just simply moving the steering wheel type of thing?

Elon Musk
CEO, Tesla

Yes. I think there'll be a transition period where people will be able to take over and should be able to take over from the Robotaxi. Then once regulators are comfortable with us not having a steering wheel, we'll just delete that and for cars that are in the fleet, and obviously with the permission of the owner, if it's owned by somebody else, we would just take the steering wheel off and put a cap where the steering wheel currently touches.

Speaker 16

There might be two phases to Robotaxi. One where the service is provided and you come in as the driver, but could potentially take over, and then in the future, there might not be a driver option. Is that how you see it as well?

Elon Musk
CEO, Tesla

In the future, the probability of the steering wheel being taken away in the future is 100%. People, consumers will demand it.

Speaker 16

Initially, you would call.

Elon Musk
CEO, Tesla

This is not I'm going to be clear. This is not me prescribing a point of view about the world. This is me predicting what consumers will demand.

Speaker 16

Yeah.

Elon Musk
CEO, Tesla

Consumers will demand in the future that people are not allowed to drive these two-ton death machines.

Speaker 16

I totally agree with that.

Elon Musk
CEO, Tesla

Yeah.

Speaker 16

In order for a Model 3 today to be part of the Robotaxi Network, when you call it, you would then get into the driver's seat, essentially-

Elon Musk
CEO, Tesla

Yeah

Speaker 16

... just to be on the same. Okay.

Elon Musk
CEO, Tesla

That's correct.

Speaker 16

That makes sense. Thank you.

Elon Musk
CEO, Tesla

Exactly. Just to sort of like, there were amphibians, pretty much things just become like land creatures. There'll be a little bit of sort of amphibian phase.

Speaker 19

Hi.

Elon Musk
CEO, Tesla

Sorry. I see with the Okay.

Speaker 19

Yes. The strategy we've heard from other players in the Robotaxi space is to select a certain municipal area to create a geo-fenced self-driving. That way you're using an HD map to have a more confined area with a bit more safety. A, we didn't hear much today around the importance of HD maps. To what extent is an HD map necessary for you?

Second, we also didn't hear much about deploying this into specific municipalities where you're working with the municipality to get the buy-in from them, and you're also getting a more defined area. What's the importance of HD maps, and to what extent are you looking at specific municipalities for rollout?

Elon Musk
CEO, Tesla

I think HD maps are a mistake. We actually had HD maps for a while. I actually canned that. You either need HD maps, in which case, if anything changes about the environment, the car will break down. You don't need HD maps, in which case, why are you wasting your time doing HD maps? The HD maps thing, the two main crutches that should not be used and will in retrospect be obviously false and foolish are LiDAR and HD maps. Mark my words.

Speaker 20

Hello.

Elon Musk
CEO, Tesla

If you need a geo-fenced area, you don't have real self-driving.

Speaker 20

Elon, it sounds like maybe battery supply could be the only bottleneck left towards this vision. Also, could you just clarify how you get the battery packs to last a million miles?

Elon Musk
CEO, Tesla

I think cells will be a constraint. That's a subject for a whole separate subject. I think we're actually going to want to push our sort of standard range plus battery more than our long-range battery, b ecause the energy content in the long-range pack is 50% higher kilowatt hours.

Essentially, you can make a third more cars, if they're all sort of standard range plus instead of the long-range pack, because one's like around 50 kWh , the other one's around 75 kWh . We're actually probably going to bias our sales intentionally towards the smaller battery pack in order to have a higher volume of. Basically, the obvious thing to do is to maximize the number of autonomous units or maximize the output that will subsequently result in the biggest autonomous fleet down the road.

We're doing a number of things in that regard, but it's just not for today's meeting.

Speaker 20

The million-mile life?

Elon Musk
CEO, Tesla

The million-mile life is basically just about getting the cycle life of the pack. You need basically on the order. Like, let's say you've got a, basic math, if you've got a 250 mi range pack, you're going to need 4,000 cycles. Very achievable. We already do that with some of our stationary storage solutions like Powerpack. We're ready to deploy Powerpack with 4,000 cycle life capability. Yeah.

Speaker 18

Can I ask?

Elon Musk
CEO, Tesla

Sorry.

Speaker 18

Yeah. I wanted to ask-

Elon Musk
CEO, Tesla

It's like ventriloquism.

Speaker 18

No, it's obviously significantly very constructive margin implications to the extent you can drive attach rates much higher over the Full Self-Driving option. I'd just be curious if you can level-set kind of where you are in terms of those attach rates and how you expect to educate consumers about the Robotaxi scenario so that attach rates do materially improve over time.

Elon Musk
CEO, Tesla

Sorry. It's a bit hard to hear your question.

Speaker 18

Yeah. Just curious where we are today in terms of Full Self-Driving attach rates, in terms of the financial implications. I think it's hugely beneficial if those attach rates materially increase because of the higher gross margin dollar that flow through to the extent people do sign up for full FSD. Just curious how you see that ramping. For what the attach rates are today versus how do you expect to educate consumers and get them aware that they should be attaching FSD to their vehicle purchases?

Elon Musk
CEO, Tesla

We're going to ramp that up massively after today. Yeah. The really fundamental message that consumers should be taking today is that it's financially insane to buy anything other than a Tesla. It will be like owning a horse in three years. Fine if you want to own a horse, but you should go into it with that expectation.

If you buy a car that does not have the hardware necessary for Full Self-Driving, it is like buying a horse. The only car that has the hardware necessary for Full Self-Driving is a Tesla. People should really think about their purchase. Any other vehicle, it's basically crazy to buy any other car than a Tesla. Yeah. We need to convey that argument clearly, and we will have to today.

Trip Chowdhry
Analyst, Global Equities Research

Perfect. Thanks for bringing the future to present.

Elon Musk
CEO, Tesla

You're welcome.

Trip Chowdhry
Analyst, Global Equities Research

Very informational time today. I was wondering, you did not talk much about Tesla pickup. Let me give a context for that. I could be wrong, but the way I'm looking at Tesla Network, as an early adopter and something as a testbed, I think Tesla's pickup may be the first phase of putting the vehicles in network because the utility of Tesla pickup would be pretty much people who are either loading a lot of stuff or are in the profession of construction or little here and there odd items, like picking up stuff from The Home Depot.

Elon Musk
CEO, Tesla

Sure.

Trip Chowdhry
Analyst, Global Equities Research

I would say that maybe it needs to have a two-stage process, pickup trucks exclusively for Tesla Network as a starting point, then people like me can buy them later. What are your thoughts on that?

Elon Musk
CEO, Tesla

Today was really just about autonomy. There's a lot that we could talk about, such as cell production, pickup truck, and future vehicles. Today was just the focus on autonomy. I agree, it's a major thing. I'm very excited for the Tesla pickup truck unveil later this year. It's going to be great.

Colin Langan
Analyst, UBS

Colin Langan in UBS. Just so that we understand the definitions, when you refer to feature complete self-driving, it sounds like you're talking level five, no geofence. Is that what's expected by the end of the year? Just so we're all.

Elon Musk
CEO, Tesla

Yes.

Colin Langan
Analyst, UBS

On the same page. The regulatory process. Have you talked to regulators about this? This seems quite an aggressive timeline from what other people have put out there. What are the hurdles that are needed, and what is the timeline to get approval? Do you need things like in California, I know they're tracking miles with an operator behind them. Do you need those things? What is that process going to look like?

Elon Musk
CEO, Tesla

Yeah. We talk to regulators around the world all the time. As we introduce additional features like Navigate on Autopilot this requires regulatory approval on a per jurisdiction basis. I think fundamentally, regulators, from my experience, are convinced by data. If you have a massive amount of data that shows that autonomy is safe, they listen to it. They may take time to digest the information. Their process may take a bit of time, but they have always come to the right conclusion from what I've seen.

Colin Langan
Analyst, UBS

Oh, I have a question over here.

Elon Musk
CEO, Tesla

I've got lights in my eyes and a pillar.

Colin Langan
Analyst, UBS

Yeah.

Elon Musk
CEO, Tesla

Okay. All right.

Colin Langan
Analyst, UBS

Some of the work we've done trying to better understand the ride-hail market, it looks like it's very concentrated in major dense urban centers. Is the way to think about this, that the Robotaxis would probably deploy more into that area and the additional Full Self-Driving for personally owned vehicles would be in the suburban areas?

Elon Musk
CEO, Tesla

Yeah. I think probably, Tesla-owned Robotaxis would be in dense urban areas along with customer vehicles. Then as you get to medium and lower density areas, it would tend to be more that people own the car and occasionally lend it out.

Trip Chowdhry
Analyst, Global Equities Research

Can you talk about some of the challenges in trying new vehicles to work in, say, Manhattan, San Francisco, and Chicago? There are so many edge cases in those places.

Elon Musk
CEO, Tesla

There are a lot of edge cases in Manhattan and, say, downtown San Francisco. There are various places around the world that have challenging urban environments. We do not expect this to be a significant issue. When I say feature complete, I mean it'll work in downtown San Francisco and downtown Manhattan this year.

Tasha Keeney
Analyst, ARK Invest

Hi, I have a neural net architecture question. Do you use different models for, say, path planning and perception, or different types of AI, and how do you split up that problem across the different pieces of autonomy?

Elon Musk
CEO, Tesla

Well, essentially, right now, AI and neural nets are used really for object recognition. We're still basically just using it as still frames, so identifying objects in still frames and tying it together in a perception path-planning layer thereafter. What's happening steadily is that the neural net is eating into the software base more and more.

Over time, we expect the neural net to do more and more. Now, from a computational cost standpoint, there are some things that are very simple for a heuristic and very difficult for a neural net. It probably makes sense to maintain some level of heuristics in the system because they're just computationally 1,000 x easier than a neural net. Like a neural net is like a cruise missile, and if you're trying to swat a fly, just use a fly swatter, not a cruise missile.

Over time, I would expect that it moves really to just training it against video, then video in, car steering and pedals out. Basically, video in, lateral and longitudinal acceleration out almost entirely. That's what we're going to use the Dojo system for. There's no system that can currently do that.

Speaker 16

Maybe over here. Just going back to the sensor suite discussion, Elon. One area I'd like to just talk about is the lack of side radars. In a situation where you have an intersection with a stop sign where there's maybe a 35, 40 mi per hour cross traffic, are you comfortable with the sensor suite, the side cameras being able to handle that? Just maybe talk a bit about that.

Elon Musk
CEO, Tesla

No problem. Essentially, the car is going to do kind of what a human would do. You can think of a human as basically a camera on a slow gimbal. It's quite remarkable that people are able to drive the car in the way that they are, because you can't look in all directions at once. The car can literally look in all directions at once with multiple cameras.

Humans are able to drive just by sort of looking this way, looking that way. They're obviously stuck in their driver's seat. They can't really get out of the driver's seat. It's like one camera on a gimbal and a conscientious driver can drive with very high safety. The cameras in the cars have a better vantage point than the person.

They're up in the B pillar or in front of the rearview mirror. They've really got a great vantage point. If you're turning onto a road that's got a lot of high-speed traffic, you can just do what a person does, just gradually turn a little bit, then go fully into the road, let the camera see what's going on, if things look good, the rear cameras don't show any oncoming traffic, off you go. If it looks sketchy, you can just pull back a little bit, just like a person. The behavior starts to become remarkably lifelike. It's quite eerie, actually. The car just starts behaving like a person.

Ben Kallo
Analyst, Baird

Over here. Here we go. Ben Kallo, right here.

Elon Musk
CEO, Tesla

Okay.

Ben Kallo
Analyst, Baird

Given all the value you're creating in your auto business by wrapping all of this technology around your cells, I guess I'm curious as to why you would still be taking some of your cell capacity and putting it into Powerwall and Powerpack. Wouldn't it make sense to put every single unit you can make into this part of your business?

Elon Musk
CEO, Tesla

We've already stolen almost all the cell lines that were meant to go to Powerwall and Powerpack, used them for Model 3. Last year, in order to make our Model 3 production and not be cell-stocked, we had to convert all of the 2170 lines at the Gigafactory to car cells. Our actual output in total gigawatt hours of stationary storage compared to vehicles is an order of magnitude different.

For stationary storage, we can basically use a whole bunch of miscellaneous cells out there. We can just gather cells from multiple suppliers all around the world. You don't have a homologation issue or safety issue like you have with cars. That's basically our stationary factory business has been just kind of feeding off scraps for quite a while. Yeah.

We really think of production as being there are many, many constraints of a massive production system. The degree to which manufacturing and supply chain is underappreciated is amazing. There are a whole series of constraints, and what is the constraint in one week may not be the constraint in another week. It's insanely difficult to make a car, especially one which is rapidly evolving. Yeah. I'll just take a few more questions, and then I think we should just break for so you can try out the cars.

Adam Jonas
Analyst, Morgan Stanley

Hi, Elon. Adam Jonas. Questions on safety. What data can you share with us today on how safe this technology is, which would obviously be important in a regulatory or insurance discussion?

Elon Musk
CEO, Tesla

Well, we publish the accidents per mile every quarter. What we see right now is that Autopilot is about twice as safe as a normal driver on average, and we expect that to increase quite a bit over time. Like I said, in the future, consumers will want to outlaw, I'm not saying they will succeed, nor am I saying I agree with this position, in the future, consumers will want to outlaw people driving their own cars because it's unsafe.

If you think of elevators used to be operated on a big lever, like you go up and down the floor, and there's like a big relay, and you had elevator operators, periodically, they would get tired or drunk or something, they'd turn the lever at the wrong time and sever somebody in half.

Now you do not have elevator operators, and it'd be quite alarming if you went into an elevator that had a big lever that could just move between floors arbitrarily. There's just buttons. In the long term, again, not a value judgment, I'm not saying I want the world to be this way. I'm saying consumers will most likely demand that people are not allowed to drive cars.

Adam Jonas
Analyst, Morgan Stanley

Elon, a follow-up. Can you share with us how much Tesla is spending on Autopilot or autonomous technology by order of magnitude on an annual basis? Thank you.

Elon Musk
CEO, Tesla

It's basically our entire expense structure. $33 billion a year.

Speaker 16

Question on the economics of the Tesla Network, just so I understand. It looked like, you get a Model 3 off lease, $25,000 goes on the balance sheet, would be an asset. Then it would cash flow at $30,000 a year, roughly.

Elon Musk
CEO, Tesla

Yeah.

Speaker 16

Is that the way to think about the economics of?

Elon Musk
CEO, Tesla

Yeah. Something like that. Yeah.

Speaker 16

Then just in terms of the financing of it, there was a question earlier, you mentioned you would do it. Is it cash flow neutral to the Robotaxi program or cash flow neutral to Tesla as a whole?

Elon Musk
CEO, Tesla

Sorry, the cash flow neutral one.

Speaker 16

He asked a question about financing the robotaxi. It looks to me like they're self-financing.

Elon Musk
CEO, Tesla

Yeah

Speaker 16

You mentioned they would be basically cash flow neutral. Is that what you were referring to?

Elon Musk
CEO, Tesla

No, I'm just saying between now and when the robotaxis are fully deployed throughout the world, the sensible thing for us is to maximize rate and drive the company to cash flow neutral.

Speaker 16

Right. Okay.

Elon Musk
CEO, Tesla

Once the Robotaxi fleet is active, I would expect to be extremely cash flow positive.

Speaker 16

You were talking about production-

Elon Musk
CEO, Tesla

Yeah.

Speaker 16

... to produce them all. Yeah. Okay. Thanks.

Elon Musk
CEO, Tesla

Maximize the number of autonomous units made.

Speaker 16

Thank you.

Elon Musk
CEO, Tesla

Okay. Just maybe one last question.

Speaker 19

Here. Hello. If I add my Tesla to the Robotaxi network, who is liable for an accident? Is it Tesla or is it me if the vehicle has an accident and harms somebody?

Elon Musk
CEO, Tesla

It's probably Tesla. I think the right thing to do is just make sure there are very, very few accidents. All right. Thanks, everyone. Please enjoy the drives. Thank you.

Adam Jonas
Analyst, Morgan Stanley

Thank you very much.