Patrick Donahue, Cloudflare
Transcript
Thank you for tuning in to today’s full episode of the Breaking Changes podcast. I’m your host and Chief Evangelist for Postman, Kin Lane. With Breaking Changes we explore specific topics from the world of APIs, but look at it through the lens of business and engineering leadership. Joining me today we have Patrick Donahue, Vice President of Product for Application Security at Cloudflare. Patrick shared Cloudflare’s vision of the API management layer that has become part of the fabric of the web, offering a very progressive view of how the API gateway fits into the business landscape, and shared their really interesting approach to treating APIs as a product. Let’s dive in with the basics. Who are you and what do you do?
Great to be here. My name is Patrick Donahue. I work for a company called Cloudflare, Vice President of Product responsible for our application security products. So that’s API security and management, web application firewall, rate limiting, all the kind of custom managed rules, DDoS, bot management, client-side security, and threat intelligence.
Nice, that’s a nice mix. I would say I’m very dependent on y’all, not just for things we do at Postman but personally. My entire DNS, increasingly my rules, my routing, all my firewall, all my encryption, and a lot of it I’m dependent on automating by your API. So I’m appreciative for all the work that you all do.
Absolutely. And I love actually going to meetings with customers. I spend a lot of my time in meetings with customers, and oftentimes, especially if they haven’t used us yet as a company, there’s people in the room that, like yourself, are using us from a personal basis. So they already know a bit about Cloudflare, and maybe we’ve stopped a DDoS attack to their site or made their DNS a little faster. So great to speak with you on that. Are you using Terraform for managing Cloudflare, or how are you actually managing your property?
Mine are actually what I consider platform ops collections. So Postman has a variety of ways where as an API client you can build little collections that do different things, and so I consider these my operational-level collections using your API. I have them dialed in, and then I have environments. I have several domains I manage through you all, different apps, and so I’ll have different workspaces and suites of collections that I use. But it is very Terraform-like. It’s just not a full tear down or build; it’s work about configuration and optimization.
And I was actually, in the early days of Cloudflare, I’ve been in the company about seven years, I was a lot closer to the code than I am now, which is a good thing for the company and the stability of our operation. But I was using Postman to do a lot of our early testing. When we build products, product managers are bus stops with us, in terms of we don’t have a QA team, which surprised me at first. And then I realized I was coming from fintech, and once I got into it I realized that makes a lot of sense, and we have millions of free customers that do a lot of QA for us. But I was using Postman to kind of move through API calls and test things. A lot of the products when we first build them, the early adopters just don’t care about the dashboard. They’re trying to do things at scale and they want to make hundreds of thousands of API calls on a daily basis to configure Cloudflare, create new properties. Take Postman for example. You have shop.postman, and that is powered by Shopify on the back end. Shopify is making API calls to us to sort of provision that hostname and get an SSL certificate for it. Those sorts of companies that have really strong engineering organizations, they don’t care about the dashboard; they want to control everything via API.
Yeah, and Postman first and foremost is known as an API client, second as a testing tool. But that automation that’s come along with testing is what’s really saved me, because I can create a collection for a specific purpose and it’s documented, and I just don’t have the memory brain cells I used to when I was younger. So I can learn part of the Cloudflare API to do a specific DNS thing that I need, dial it all in, document it as a collection, it runs. I have the domain and the zone stuff kind of abstracted away so I can use it across different zones and different domains, and then that’s there in a workspace. And I forget that knowledge and I go back about my business, I come back and it’s all documented, and I can have that scheduled, I can have it as part of a pipeline. So it’s really just my memory about how things work, because I can’t keep up with what you all are building sometimes, because we need to keep our velocity high.
Yeah, and we like to be shipping. That’s fun. At the end of the day as a product person you want to put new things out there and solve customer problems. Speaking of your collections for documentation, I remember using, is it Newman, your CLI? We would write collections and run them in Docker containers to run through full end-to-end tests, you know, making a new release and making sure nothing regressed or things like that.
Yeah, and I do that. I have a couple domains that are real heavy sub-domain. There’s literally a couple hundred sub-domains and it’s a very distributed architecture, but I have collections that will audit that and use the Cloudflare API to get the latest configurations, kind of test what’s going on and make sure it’s all set up and working. And it’s pretty robust and automated in the way that it helps me keep tabs of a pretty sprawling distributed landscape within a domain, and it helps me automate across that and ensure quality and governance and stuff like that.
So that’s great. One of the things we tell our teams is if you’re in the dashboard you can see, there’s a button you can click but there’s a little drawer beneath it, and you can click that and it’ll show you the underlying API call. So we try to make it as easy as possible to do that. In a lot of cases, and some of the new stuff we’re building, it’ll interpolate what you’ve typed in into the API file so that you can actually just copy and paste it and not worry about it.
I’ve written about that probably two or three times, and I’ve referenced it in probably about 30 or 40 conversations or talks that I’ve given, that specific API call. For me that is how UI should be. It’s like, give me the API behind every UI action that I need and put it right there for me. So you guys are really that poster child of doing it well in my book.
That’s great to hear. What makes it easier is that a lot of the companies that I talk to, especially those adopting our API management solution which we recently announced, their front ends are essentially just lightweight wrappers around the APIs that their customers are calling to manage their properties, that don’t want to use the dashboard. That’s how we build as well, so it makes it relatively easy to do that, when you’re just wrapping the APIs in the dashboard.
Yeah, no, it’s the way I see the world, but I’m API biased obviously. But you kind of brought it home. I wanted to jump on this call with you and have a conversation. For me Cloudflare is just, I get DNS out of, I spent decades in DNS hell, and the way that you guys have all simplified DNS for me, I’ll never go back to that reality. Like, what did you do?
Oh yeah. I had my own, I had multiple C blocks back in the 90s and 2000s. I was in the weeds, in the days where you make a DNS screw-up and you have to live with it for a while until things resolve.
Yeah, and I never want to go back. So once you guys laid that foundation I was like, I just fully respect what you’re doing. But then making the encryption piece is another. You’re like, I get certs but I don’t want to live in that world, I want to do what I do. And then you guys have been doing this incrementally. Some of the edge, the runners, how you guys perceive serverless, that maybe is another show. But you guys jumped on the API management train recently and I want to hear why. What’s the story behind it? How do you guys see all of this?
Yeah, absolutely. Just laughing at your utilization of those flash 24s that you had. These days those are probably worth a lot, faster appreciating than Bitcoin, so I don’t know if you still have them, but they’re quite expensive these days. I administered buying four and eight when I was the student manager of the network operations center in college, so I know the pain. We had a very flat slash 16 for the entire campus. But no, I know the pain managing DNS. There’s a great t-shirt that says, you know, it was DNS. It’s always DNS when something goes down.
So just to give a little background before I jump into why we entered this space. Some of the things you mentioned, issuing SSL/TLS certificates, managing DNS, those are things that just have to be done, and have to be done well and reliably and securely, and that’s not really core to what most people do, unless you’re running an infrastructure company and providing that for others. That’s something you just have to do to run your business. So we try to take care of as much of that as possible to let you focus on building and growing your business, the business logic and the things that only you can do well, versus things that we can kind of take off your plate. We like to think about the company as helping build a better internet. We’ve essentially built this massive global network that spans the world. I think we’re in 250 cities and over 100 countries now, and we’re very close to the eyeballs, about 50 milliseconds from 95% of the internet-connected population. So what we try to do, to get to the API side, is we want companies and developers to be able to just plug into that network and automatically improve the security of everything they do online. And also the way we do it makes it more performant and reliable. We started out focusing on security, and Matthew, our CEO, likes to tell a story of customers would write in in the early days and say, hey, this is actually much faster now, but that wasn’t the original focus. We do do a lot on performance and we get a lot more into the networking side these days, but that wasn’t part of the original outset for the company.
So we’re on a mission to get to APIs. We’re on a mission to go through all the boxes that you and I probably back in the late 90s, early 2000s got into a data center and racked and stacked. I enumerated some of them, the web application firewall, the load balancer, DDoS scrubbing boxes, I could go on and on. But one of the ones that people put a lot of logic in is some sort of load balancer or gateway that does things like TLS termination, does routing, does some sort of security function, rate limiting, et cetera. If you have one of those running in a data center and everything is there and you don’t care about redundancy or being close to users around the world, that probably works okay. But a lot of our customers are running either multi-region in the same cloud provider or multiple cloud providers, and they use us at the edge to attract all their traffic to those pops that I mentioned, and then perform as much of those operations you would typically do downstream in a data center at the edge. So we’ve essentially taken all those boxes and written software to virtualize them, and these days we’re replacing, if you’re doing MPLS, I used to deploy that back in the day, you can do that now in software functions we’ve written into the stack at the edge.
With APIs, especially for customers that are already putting traffic through us for reasons of protecting against DDoS or volumetric attacks, or even caching certain assets, that traffic is already passing through us. We’ve had customers come to us and say, hey. What we’d like initially was, we’d like you to do more security functions. Things like sequential abuse detection, we can get into that, I think it’s a really interesting topic. Or, we want to write rules that use the intelligence that you’re gleaning from those millions of sites that are running through you and apply that and take different action at the edge. But more recently it’s customers coming to us and saying, hey, we’re spending hundreds of thousands of dollars on this API gateway solution and we’re only using a very small fraction of it. We’d love, given the traffic’s already running through you, for you to implement those remaining functions that we are using, so that we can simplify and consolidate from a vendor perspective and just deal with that at the edge and not have to worry about it downstream.
Yeah, well, there’s a lot to unpack there. I’ve been following the API management game, I started my blog API Evangelist in 2010 to understand the cloud and mobile. But at the time 3scale, Apigee, MuleSoft, all of their evolution, the notion was, here’s a whole suite of features that you need, a portal, docs, all of these things. I consider it now our grandfather’s gateway. It’s just a massive trunk of everything that you don’t need but you kind of accumulated over the years. And then you see this next generation of Kong and Tyk and others emerge. They’re like, hey, we need a smaller, lightweight gateway. But then as these things have been commoditized, I would say around 2015 things started being commoditized at the gateway layer, some of those features are just essential. The ones you talked about, routing, rate limiting, security, caching, these things are fundamental, core, but they’re commodities now. They just need to be baked in, they need to be there. And I kind of feel like that’s reflected of Cloudflare’s approach to things. Once things reach that essential-grade need, you guys abstract and simplify it, and it just becomes part of the fabric of the web.
Yeah, that’s definitely true for a lot of the functionality. We’re building it into that engine at the edge, which matches on requests, and we see about 32 million requests per second, so a lot of stuff to match on. And then, what action do you want to take on that, and how much information can we give you about the request itself, the session, the user behind it, to decide what you want to do. Those mature functions are quite simple to implement. I think where a lot of our focus now is, is what are the unique insights that we can surface that you might not be aware of. If you’re running a mobile application, that’s calling APIs at the end of the day, calling APIs in the back end to perform certain operations. Some of the really cool stuff that we’re working on, that I’m really excited for to go GA for customers, and we’ve done some manual tests and shared some data, are things that are looking at the patterns of how people are using APIs and trying to surface for them what is anomalous.
Take a food delivery, this is a real-world example, I won’t mention the customer by name. What we were doing with them is we said, okay, let’s take all that traffic passing through us, let’s see how many unique different API calls there are, and there was upwards of a hundred thousand, because you’ve got unique identifiers and things like that that are part of the path. Can we normalize that down to remove those super-unique identifiers and put some variable placeholders there? I think it was 60, 70 API paths ultimately that got called by their mobile app. And let’s surface what doesn’t look right that your security teams might want to dig in on. If you’re running an operation at scale, you have so much data to sift through. Where do you spend your time, how do you triage? So what we did there is we built these Markov chains that said, what is the probability to move from state one to state two to state three to state four, all the states that we identified. If you’re using a food delivery app, you’ll open the app and you’ll log in, you’ll see a list of restaurants that are delivering to your home, you might scroll through, pick one, see the items, add a few. And then if you’re like me, you change your mind and you go to a different restaurant, clear your cart, add some other items, and eventually place an order. That’s the behavior that most people will navigate this API through. They’re not really aware they’re calling an API, but their mobile application is. What you wouldn’t do is log in and then enumerate every single restaurant and every single price in a very short period of time, or even just do that regardless of the amount of time. In this particular case, the customer’s hypothesis was that this was actually a competitor running their mobile app in an emulator, to try to get around some of the security controls. We were able to surface that insight for them by running this against our Markov chain models. This is something that we’re working to automate now. That’s where I’m really excited. The basic stuff you mentioned that’s commoditized, anybody can do that. The harder stuff that’s more useful and reduces time for security teams and development teams is where we want to focus.
And just to give you one example on the development side, that was a security example. We’ve got this product, Cloudflare Workers, a serverless product where you can run code around the world. We have the concept, we like to say the region is global, where you don’t pick anything, you just write some code, you push it to us, and it runs on every machine and every data center around the world. Those paths initially were, can you modify the functionality of what Cloudflare does? Cloudflare doesn’t have a button to do this, can we tweak it to do that based on writing JavaScript and using the V8 isolates that Google and Chrome developed? That was the initial use case. But today most people coming to this from a developer perspective are writing their applications, they’re shifting from running them in AWS in a handful of data centers to running them at the edge, and they’re implementing APIs there. So how can we, if we’re seeing this traffic pass through us, say, hey, this is something that maybe we could run much quicker for you at the edge, and maybe generate some boilerplate code or some scaffolding that you could modify to hook into a request and a response object. So some of that more intelligent stuff is what I’m really excited about.
I think that reflects what I’m seeing. The last grandfather’s generation of API management was SOA, service-oriented architecture, web services, very top-down, heavy. I consider the second wave about these RESTful resources that we needed, images, videos, messages. But it’s very CRUD, you take your database and you create, read, update, delete; it’s very resource-driven. But what I’m seeing the next generation of APIs is very capability or, as you say, behavior based. These are the series of API calls against many resources, internal, partner, and external, that we need. These capabilities are what we’re putting out there, and then the resulting behaviors on top of that. That’s huge. That’s a goldmine of data for me as an API producer, to stay in tune with the feedback loops I need from my consumers.
Yeah, that feedback loop is a great point, and that’s how we think about how we surface insights. We launched something that we call the Cloudflare Security Center, not the most unique name, there’s a lot of security centers out there, but it shows you insights into where you should be paying attention. Basic things like, is there a dangling A record. But we’re really tightening that feedback loop for APIs, saying, hey, we’re calculating the quantiles, the p50, the p90, the p99 request rate, and then we see these outliers for you on this volumetric detection, click here to deploy a rule that already includes those thresholds that we’ve calculated for you. That feedback loop of insights, that by us seeing the traffic flow we can make easier for you, so you don’t have to figure out, those are the things that, we want developers writing code and building their applications. We don’t want them worrying about how do you expose and protect and log and audit them. As much as possible we have these conversations with the developers and the DevOps crews managing these APIs and saying, what are these things we can take off your plate? Beyond the basic stuff we talked about, a lot of our roadmap is, what can we take off your plate and let you focus on building your business?
Yeah, keep them focused on the value they develop against. That’s like what I was saying with DNS and encryption. I know this stuff and I can live in this world, but man, I need to be building other things. So the global platform that you all have built, I’m assuming the reason why people want availability in all these regions is performance. But are you seeing any sort of regulatory data sovereignty or other motivations creep into this?
Absolutely, yeah. The regions you think about, and actually to go back to the SSL/TLS point, we did this originally for encryption. We had customers tell us, hey, I only want my traffic decryptable in the US, or I only want it outside the US, or in the European Union, or actually these specific sets of data centers of your 250-plus cities. So the first thing we built was, we can build this thing that we call Keyless SSL, which is kind of a misnomer, obviously there’s a private key somewhere, but where that private key is held is controllable by the customer. We can route traffic to get to particular data centers or regions to decrypt it. That’s expanded over the years, we have something like a data localization suite. This particularly comes up, about half of our business is outside the US as a global company. I spent a couple years in our London office and would meet with customers regularly, especially in Germany, that would say, hey, I want this traffic not to go to certain parts of the world, and I understand it may be slower if you’re accessing it from Australia or from the US, but I care more about data sovereignty and related concerns. We can do a lot of cool stuff because we’re running essentially a programmable network, and we’re writing all this stuff in the Linux kernel, eXpress Data Path, eBPF functions. We can program our network in a way that others can’t. We’re able to say, attract traffic only to particular IPs or IP prefixes in certain parts of the world.
By default, if you put something on Cloudflare, your site for example will announce those, somebody does a DNS resolution request and the IP returned is the same IP wherever you are in the world. We use IP anycast technology, we’re announcing those prefixes to our peers. I’m in Austin, Texas right now, so if I do a resolution for your site, it’s going to hand back an IP and my ISP is going to route me to our Dallas data center, probably a single hop away in terms of where we’re peered directly, Google and my ISP and Cloudflare are probably peered in that data center. But if that data center falls out, maybe we’re doing maintenance, we’re shutting it down, Google is going to see my ISP, going to see what’s the next hop I can get to. That’s the default behavior. But sometimes people want to say, only serve these IP addresses or announce them from particular regions and countries. That’s a lot of the power for Workers, for example. We thought initially the hypothesis was, you can move your code to the edge and it’s going to be that much faster, and if you’re 10, 20 milliseconds away from somebody you can do all these real-time things you wouldn’t be able to do otherwise. And actually a lot of the use cases and where we’re seeing a lot of the demand is really on that data localization side, adhering to the changing regulatory framework and saying, I only want to process this here, and we know you have data centers in these countries that we need to keep the traffic and only process them there. When you’re in product management you have these ideas in your head, when you write the blog post, what people are going to use it for, and then they use it for something totally different and it surprises you. That’s a cool experience, a rewarding experience.
Yeah, I did a show a week or so ago with the CEO of Open Collective. They’re the funding model behind a lot of open source projects, and they’re really building in a lot of the other scaffolding that open source projects need to run a business, to do invoices in different countries and operate as open source globally. I’ve always felt like that’s what Cloudflare has been doing for me, just abstracting away a lot of these global concerns that I have, I just don’t have the time for. I’m a small shop, small business, busy in my enterprise, and so you all are abstracting away, like I said, the DNS, the encryption, but then these other regulatory, I just don’t have the time. I get European Union, but Germany versus France versus UK, I don’t have time to think about all these things, but it changes whether you’re in the EU or not.
Right, and so we have a phenomenal public policy team that is global, and they’re just constantly monitoring this. Again, the goal is to monitor it and pick it off your hands so that you don’t have to do it. The cool thing about the localization we’re doing, a lot of these ideas come initially from our team internally. We will meet with different teams and they say, hey, we’re using this third-party solution today, or something we kind of cobbled together internally, but this might be a good product for our customers to use, and we’ve got some seeds of it, and we think we should maybe build it in a way that others can use. That’s where a lot of the ideas come from. One example is the API discovery that we’re doing. I sat down with our CSO and tried to get some ideas, and he said, hey, when I was at Uber, the first thing I was concerned about when I joined is, what are all the APIs out there that my developers are publishing that I might not be aware of, to the world. So we said, great, let’s automatically catalog those for you so that you can apply security policies to them and be aware of performance and what data they’re returning. And then we also had internal folks say, hey, I’m worried that my APIs might return something I don’t want them to. So then we built the ability to inspect that traffic as it goes back to the caller and say, are there certain strings or patterns you want to match for and block?
Another example there, our dashboard, people try to credential stuff the dashboard. They’ll download, you probably are on Have I Been Pwned and you probably got emails from Troy Hunt before if you’re subscribed there, some company leaked credentials you had, and people will gather those up and start running those and trying to hit our APIs in our dashboard. So what the team that manages that internally said is, hey, can you help us, they may be flying under a rate limit and not tripping that, but can you help us identify when someone’s trying to log in with credentials that have been exposed somewhere on the internet? So we have a way where customers can opt in and say, hey, if somebody’s using their credential that’s dropped in one of these databases, either just block it right from that API authentication call, or, what a lot of people want to do, they want to add a request header so that when it goes downstream to their origin server they can decide what to do. Maybe they’ll serve an additional form of authentication, maybe they’ll send a password reset. So we try to give you those signals in the API call flow to take additional action downstream if you want, or at the edge.
Yeah, one of the first things I stumbled on with Cloudflare that tripped me up but then now is a bedrock cornerstone is the analytics. When I was first running all of my applications and websites and properties through Cloudflare, I had Google Analytics running and I was very much in a Google mindset. And then you switch to the Cloudflare analytics, you’re looking at me, and it tells a different story. At first I was like, hey, this just seems wrong, and I even had some conversations with your teams about it, and they’re like, well, it’s a different view of the landscape. Over time I realized how beholden I’d become to the Google performance of analytics, that’s very advertising-driven. It was like frog cooking in a pot over time, with every change I was kind of more and more beholden. And then Cloudflare has given me this entirely honest view of the network, and I realized, I don’t care about page views in a vanity, ego kind of way. I care about my properties and what’s going on, and I’m definitely a technologist, I have a different view. I don’t look at Google Analytics anymore. So I can see this when it comes to the API management layer, like I feel like my API management analytics have been lying to me, or lying’s a strong word, but leading me down a specific narrative.
Yeah, that’s a good point. Just to spend a second on analytics, we have all this data, and in the last couple years, to your point, we’ve been investing heavily and sharing out this data to give you the insights you need, but doing it in a privacy-preserving way. Some of the critiques you’ll see with things like Google Analytics and their reCAPTCHA service, for example, is Google’s just gathering and holding all this data so that they can sell you advertising. Whenever we try to build a product, this is kind of a core principle of Cloudflare, we think, how can we help the end user become more secure and private, in particular in how they’re actually accessing this. The difference there, of course, is a lot of the browsers will block, I use Brave, a Chromium version that’s kind of privacy-focused, it’ll block a lot of those by default, outbound calls to third parties that are capturing your data. The difference is, in this case the data is hitting Cloudflare for requests to land there, so we’re in a different vantage point to be able to gather it than some JavaScript that was added to the page. The other thing is that JavaScript that’s serving in a page is a security risk to you. You’ve probably heard of like Magecart-style attacks, where somebody gets access to, or through a supply chain attack gets access to some JavaScript running in your browser, makes a change, and all of a sudden the data you’re sending, your credit card number, is going to the attacker’s server versus where you intend it to go. So we’re really trying to focus on reducing the JavaScript that’s actually served in the browser and monitoring for changes and letting people know. Anyway, we’re getting a little out of the API world.
Yeah, that’s a really neat part of JavaScript surface reduction there. No, it makes a lot of sense. And I feel like the complexity of the systems that we’ve built, both internally, partner, and external, the SaaS services we’re using, everybody I go to and talk to, I go into enterprise organizations, I went into the IRS and did an all-day workshop, no one knows where all their APIs are. Nobody I talk to can tell me where they are. And I go, well, do you have an API catalog, and they’re like, yeah, we’ve got IBM this or we’ve got Apigee this, but it’s not up to date.
Right, we’re just moving too fast.
Exactly, and we can’t get teams to update the metadata and it’s just out of date. So our documentation suffers, our security suffers, our performance, all these things suffer because of this. And us in the API space say, well, you’ve got to go design-first and build this contract and then everything downstream will magically work, and that doesn’t work either. We’re just moving fast and very often breaking things, and we’re just responding to business needs and customers. So do you feel like that automation and the intelligence that you talked about is the only way we’re going to be able to manage this chaos and get it to any sort of state that suits our business needs?
Yeah, automation, if you don’t automate something it’s going to follow the manual effort. We think about that internally, just for processes unrelated to this. If there’s a human involved in onboarding a customer, to take that network product you mentioned earlier, initially if you want to take your IP space, imagine you still own those class blocks and you wanted to announce them at our edge and attract traffic there, that’s a project called Magic Transit, if you wanted to do that early on it was a bunch of humans involved, and potentially there could be steps missed. You’d have to talk to the networking team to prepare the routers, you’d have to talk to the SRE team to do various operations on what we call metals, all the servers in the data centers, and steps would get missed. So we had a concerted effort, the product manager and that team, to automate that stuff. It’s not 100% automated, but it’s pretty close. I see the same thing in talking to customers that are building and running APIs. If you don’t automate it, things are going to get missed, and that’s just human nature. We’re not machines, we’re going to make mistakes, we have time commitments and crunches. So what are those things that we can automate for you? One good example is a lot of our customers like to validate the schema of a request using OpenAPI format when a request comes in. Some of them have great tooling that their system will automatically update those schema files and push them to our edge. Others don’t have great tooling yet, they want to get there eventually, but it’s not something they have just yet. So how can we try to automatically do that for them? A lot of the ways we build product is we give you the controls and the levers initially if you want to solve it yourself with uploading your schema file, and then we think, okay, how can we do better than that? Can we automatically identify, hey, it looks like this API contract has changed, can we just let you confirm it and have us maintain that schema file for you? Those are the types of things that we think about in the API world, what are those other manual steps that you would typically take, and how can we take those off and automate them for you.
Yeah, huge, because that rate of change reflects the reality on the ground and I just can’t keep up with it. And that’s my own system, my own APIs. But you’re all also bringing in your wider security knowledge and view that you guys have of the entire landscape and building that into the intelligence and automation at that layer as well.
Yeah, and things are getting harder, not easier, on the security front, and there’s this trade-off between privacy and security which is really interesting. You may have seen the Apple iCloud Plus Private Relay announcement, where your phone is now sending through a series of proxies to the end server that’s running the API. It’s getting a lot of requests from the same API, or a small pool of APIs, that used to be individualized identifiers. You used to be able to say, okay, this IP is doing something suspicious and maybe we should block or challenge it, and this IP we’ve only seen good traffic from. But all of this stuff is blended together in the name of privacy, which we’re huge fans of and we support wherever we can, but that makes it harder to disambiguate the color of the API and perform security actions like rate limiting, for example. So there’s new approaches that need to be taken to advance those security efforts, and we’re spending a lot of our time on open protocols that will help with this. Things like device attestation and other sorts of signals that can be sent in lieu of this IP as a reputation. There’s open standards for a lot of this stuff, and if you had all the time in the world it might be a fun project to go implement yourself. But again, you don’t have that time, and so we want to make that, our CEO and our CTO constantly push us to make it magical, like make it just work. In an ideal world the user doesn’t have to do anything. That’s the SSL certificate issuance that you talked about before. As soon as we’re authoritative for your domain, we can make an API call to a certificate authority, ask what we need to do to demonstrate control, those DCV records, issue a certificate or multiple, typically, and push it to the edge. So that’s something you don’t have to do anything for. We want to do the same things for APIs and other aspects of our product. That use case where you’re seeing a bunch of clients coming from fewer IPs, can we automatically make our rate limiting work so that it doesn’t need to rely on IP addresses, they can look more to the session? Those are areas of focus for us these days.
Okay, it says a lot that I still remember my IP block. 207.189.150.32 to 154, I still, after 20 years, typing those IPs so much.
Yeah, I know the feeling. 139.140 slash 16 was the one I used to administer.
Wow, yeah, all the trauma and PTSD from those days. Cloudflare has saved me so much pain and suffering, and so I’m really impressed and fascinated with what y’all have done at the DNS layer, because back in the day, DNS is an important link in the chain, but you guys took it to a whole other level. All the edge stuff, the serverless view, I just never imagined this being such a rich layer of our reality, security. So when it comes to prioritization at this API gateway level, security obviously is probably top, performance. When it comes to customer needs, what are those top requests or needs?
Yeah, absolutely. This was kind of the fun part of being in product. The way I like to tell it internally is, some people think externally you’re having a thousand different conversations with prospects and customers, you’re talking to a bunch of different analysts, and you have this perfectly ordered list of priorities that in an ideal world will solve everything for everyone, and then you march through it sequentially. That’s not really how it works in the real world for us. What I love about us is we’ll put out announcements very early and we’ll release products early in the life cycle of the product. When I joined seven years ago that was a little uncomfortable for me, I was coming from a financial technology world where you wouldn’t ship something until it was 100% fully complete and feature-rich. We like to give stuff very early in the process to customers because we want them kicking the tires and telling us, you implemented it this way, but we actually think it would be better that way. And then you have a number of those conversations and you say, you know what, they’re right. We had this idea in these planning meetings and when customers started using it it was different. That’s how we approach a lot of the product prioritization. We’ll make an announcement, we’ll have some features customers can use, and then we’ll do a bunch of discovery calls where we’re getting on the phone and asking customers, what problems are you trying to solve, and writing those down. To get to your point of what actually people care about, a lot of it is just the discovery process of first finding what’s out there and then guiding them through that process. You might be implementing just the 20 or 30% initially that really matters to people, that they’re willing to pay for, and the other stuff they’re not even using, to your grandfather’s trunk point earlier. You’re building it in a way that’s really easy to use, especially if your traffic is already running through Cloudflare. So we’re trying to say, what are the strengths we can play to, to get people using our products and getting feedback on it. So discovery is one of them.
And then on the security topic, we like to think about reducing the noise as much as possible to only send the request to your origin that you want to see. The first layer of that, we think of it like a funnel, is just get rid of the network-layer DDoS stuff. If you’re only exposing TCP 443 on your origin, or even not exposing anything and using our Tunnel product to get it there, you don’t want to see traffic hitting that API endpoint coming from UDP traffic or TCP on different ports and protocols. So the focus on the security side is progressively reducing the noise. Volumetric DDoS, rate limiting, schema validation, how can we only send to you what really matters? And then a lot of the other needs, or the features, are opportunistic. Within our WAF product we’ve been working on the concept of managed lists, which are lists that we maintain. We’ll go out and scan the internet, find all the open SOCKS proxies that are out there, or the VPN servers, or Tor exit nodes, malware C2 servers, and all that stuff, and we can maintain that list for you. Within your API gateway config you can say, hey, if a request is coming from behind a VPN, I want to treat it differently. Or if a request is coming from an open residential proxy, which typically is indicative of a compromised device, we saw a massive attack not too long ago from these MikroTik routers that got compromised and were able to be used in an HTTP proxy way, we got hit with this, or one of our customers did, with a 17-million-request-per-second attack, so how do we block that using intelligence and just filtering the noise out?
From a priority perspective, just to go back to the discovery point, we have these conversations with customers, we implement what we need to get them successful and up and running, and we like early adopters. Postman, for example, would be a terrific, we want really smart engineering organizations that we can partner closely with, and they like to think of us as an extension of their own internal development teams, and we love those early adopters because they help us build great products. So that’s where we’re focused right now. The other things are taking those routing engines and response engines and all the stuff we’ve built into what we call Transform Rules and packaging them up in that experience to go from discovery to having a policy. And then the last piece where we’ve heard a lot of requests is, how do you know if your APIs are healthy? How do you actually know, are they under too much load, did I push a change that makes it a much more expensive call than it used to be? I’d love to get your perspective here too, I use this as part of a little discovery call. What actually means an API is struggling? Is it a p90 response time spiking by X percent for that longer tail of slow calls, or something else? Is it a rate of 5xx errors that are coming back or spiking? Those are the things, again, pointing out where you should spend your limited time, that we’re focused on right now.
Yeah, wow. First off, I love your product-management-centric approach. This is the top thing that we’re working with large enterprise customers on, is that journey you describe. They’re super risk-concerned, and they don’t understand the confidence that comes with a product management mindset, with putting something out there having those feedback loops in place. So API product management, treating your APIs as a product, is the top area we’re working with large customers on, and part of that feedback loop is directly with your customers, having that feedback loop and having that confidence that you can put things out there and get the feedback you need. But the analytics, the money, is this something they’ll pay for, what’s the volume of usage. It sounds like you guys have a very healthy view of this, but it sounds like you’re also going to extend that as part of your gateway so that others can put out APIs confidently early on, get feedback, iterate fast, move quickly in the way they need to, and have, it’s up to them to have the feedback loop with their customers, but they have the analytics, the awareness, the intelligence that they need to make the right decisions.
Yeah, that’s a great point, putting it out and letting them quickly get that feedback and not have to build in a lot of that telemetry and instrumentation, just getting the logic running and letting us tell you, okay, here’s what we see happening from our vantage point. The other thing I’ll mention is, yes, we serve a lot of very large enterprises, but a lot of the really fun customers to work with are the ones that are either paying us nothing or twenty dollars a month or two hundred dollars a month. I interact with a lot of them on Twitter, and I pay you for my personal stuff, we love it. A lot of the richest feedback we get is from those what we call self-serve or pay-as-you-go customers that aren’t associated with a salesperson or a customer success person. They write some of the most useful feedback. That’s the beauty of having this freemium model, you have this long tail of millions and millions of sites, and people are doing things that are unique. We actually see attacks against some of the smallest customers that are some of the most sophisticated attacks, and then we can learn from those attacks and put that back into the product, typically in an automated fashion, training machine learning models to protect the biggest customers. When we’re launching a new product or getting ready, I’ll just send a tweet out saying, hey, anybody interested in giving feedback on XYZ or wanting to get beta access to something, and there’s a really rich community that is excited for new products and is always willing to share their feedback. If you’re involved in those early conversations with our product team, you have an opportunity to shape how the actual product gets delivered in a way that is more useful for you. So I’d love anybody listening to this, if you see us putting out new products, reach out on Twitter, drop an email, we’re always happy to take feedback from all over.
Well, so I run our Open Tech team, which is our DevRel, developer relations, about 10 of us. Then I have what’s called our platform team, where they focus on life cycle and governance. And then I have an Open Tech team where they focus on specifications, Swagger, OpenAPI, AsyncAPI, and the standards and tooling around those. They’re going to be playing around hacking. The meeting I had before this was where we’re kind of writing a state of the gateway reality, like where are gateways right now. They’ve been commoditized for about six, seven years, who’s the current list of players, what’s the future look like, and what you’ve all just released is going to be a big part of that. So I’ll have lots of feedback, lots of thoughts from our team, we’ll connect offline. But I know you guys are just getting going, what does the roadmap look like? What can you share about where you’re headed next?
Yeah, so we’ve had a security-focused API product for a while, and so a lot of what we’re doing now is the management side of it. Some of the stuff I talked about, those routing engines and transformations and being able to get richer analytics. Other areas we’re focused on are, what are the other protocols that people care about? Most of the traffic is RESTful traffic, but we do have people coming to us and saying, my clients and partners are all calling these RESTful at the edge, but I want to turn around and transform that, bridge that to gRPC or GraphQL or other protocols that people are looking to implement. When we launch features we pick what is the largest use case to begin with to get people working on it, and then someone will come to us and say, here’s what I’m trying to do, I want to move to gRPC binary format calls for most of my traffic, and they have a date for it. That’s hard to do if you’re just on the origin, you’re going to spin up things and manage different configuration. But if we can do that transformation, go back to the Workers point in serverless, if we can give you those primitives within a programmable environment, you can do a lot with that without having to change your code downstream, or maybe you route it to a different endpoint that’s accepting gRPC or GraphQL in lieu of these RESTful calls. So that’s a big area for us. And then the invocation of Workers, putting Workers and the gateway close together, I think this is going to be one of our superpowers. If you’re going to call slash foo on an API, you can send that downstream, but if that’s logic, especially as we build out our storage solution at the edge where you have richer database capabilities on our network, you can move more of that logic to the edge and be kind of cloud-agnostic. Tightening that feedback, where you can take a path and say, okay, run this Worker, or even a more lightweight version of it, run this function, and give an interface to make that easy. That’s where a lot of the effort is going to be going, bringing Workers and API gateway very close together, and that’s something I’m really excited about. That team’s velocity of what they’re shipping is insane, and anytime we can tie things together like that from a product perspective it opens up a lot more use cases for people to implement.
Yeah, so exciting. Looking forward to working and staying in tune with this, where we just went from a synchronous API client, HTTP, to we’re now doing WebSockets and gRPC and trying to weave that into the whole Postman, so you have everything you know about core Postman but now in these other realms. So let’s keep comparing notes there. I think the biggest failure of API management over the last decade is addressing API deployment. Most people ask me, which API management provider should I use to deploy my APIs, and I’m like, that’s really a tough question, because until recently most of them wouldn’t, and so that question’s always just kind of been left hanging. Serverless I think is potentially reinventing that, but I think y’all’s approach, that intersection of at the edge, is really important to the future of what is API deployment combined with management combined with all of this multi-protocol, multi-cloud approach. So looking forward.
Yeah, I’m excited, and I would love your feedback. Get you access to the beta and your team, and would love to hear any thoughts you have. You’ve got a direct line for future requests, so let me know what we can add, we’ll be in touch. We’re big fans.
Wow, well, we went way longer than I think most of these shows go, we went almost the full hour. But I appreciate your time today, Patrick, this has been awesome.
Yeah, this has been a fun conversation, really enjoyed it. Thanks for having me on, would love to chat again in the future.
Thanks again to Patrick for stopping by. You can find more about Cloudflare at cloudflare.com, and you can find Patrick on LinkedIn. You can subscribe to the Breaking Changes podcast at postman.com/events/breaking-changes. I’m your host, Kin Lane, and until next time, cheers.
