<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:media="http://search.yahoo.com/mrss/"><channel><title><![CDATA[The Cloudflare Blog]]></title><description><![CDATA[Get the latest news on how products at Cloudflare are built, technologies used, and join the teams helping to build a better Internet.]]></description><link>https://blog.cloudflare.com/</link><image><url>http://blog.cloudflare.com/favicon.png</url><title>The Cloudflare Blog</title><link>https://blog.cloudflare.com/</link></image><generator>Ghost 3.5</generator><lastBuildDate>Tue, 06 Sep 2022 20:08:10 GMT</lastBuildDate><atom:link href="https://blog.cloudflare.com/rss/" rel="self" type="application/rss+xml"/><ttl>60</ttl><item><title><![CDATA[Cloudflare named a Leader by Gartner]]></title><description><![CDATA[Gartner has recognised Cloudflare as a Leader in the 2022 "Gartner® Magic Quadrant™ for Web Application and API Protection (WAAP)" report that evaluated 11 vendors for their ‘ability to execute’ and ‘completeness of vision’]]></description><link>https://blog.cloudflare.com/cloudflare-waap-named-leader-gartner-magic-quadrant-2022/</link><guid isPermaLink="false">63160bfeec6c4e000bf2b128</guid><category><![CDATA[WAF]]></category><category><![CDATA[Security]]></category><category><![CDATA[API Security]]></category><category><![CDATA[DDoS]]></category><category><![CDATA[Bot Management]]></category><category><![CDATA[Page Shield]]></category><category><![CDATA[Gartner]]></category><dc:creator><![CDATA[Michael Tremante]]></dc:creator><pubDate>Tue, 06 Sep 2022 16:15:44 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/09/image1-1-2-1.png" medium="image"/><content:encoded><![CDATA[<figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/BDES-3702-Gartner-MQ-Social_Blue_V2_1200x628_NOCTA--1-.png" class="kg-image" alt="Cloudflare named a Leader by Gartner"></figure><img src="http://blog.cloudflare.com/content/images/2022/09/image1-1-2-1.png" alt="Cloudflare named a Leader by Gartner"><p>Gartner has recognised Cloudflare as a Leader in the 2022 "Gartner® Magic Quadrant™ for Web Application and API Protection (WAAP)" report that evaluated 11 vendors for their ‘ability to execute’ and ‘completeness of vision’. </p><p><strong>You can register for a complimentary copy of the report <a href="https://www.cloudflare.com/lp/gartner-magic-quadrant-waap-2022/">here</a>.</strong></p><p>We believe this achievement highlights our continued commitment and investment in this space as we aim to provide better and more effective security solutions to our users and customers.</p><h2 id="keeping-up-with-application-security">Keeping up with application security</h2><p>With over 36 million HTTP requests per second being processed by the Cloudflare global network we get unprecedented visibility into network patterns and attack vectors. This scale allows us to effectively differentiate clean traffic from malicious, resulting in about <a href="http://blog.cloudflare.com/application-security/">1 in every 10 HTTP requests proxied by Cloudflare being mitigated at the edge</a> by our WAAP portfolio.</p><p>Visibility is not enough, and as new use cases and patterns emerge, we invest in research and new product development. For example, <a href="http://blog.cloudflare.com/landscape-of-api-traffic/">API traffic is increasing</a> (55%+ of total traffic) and we don’t expect this trend to slow down. To help customers with these new workloads, our <a href="http://blog.cloudflare.com/api-gateway/">API Gateway</a> builds upon our <a href="https://www.cloudflare.com/waf/">WAF</a> to provide better visibility and mitigations for well-structured API traffic for which we’ve observed different attack profiles compared to standard web based applications.</p><p>We believe our continued investment in application security has helped us gain our position in this space, and we’d like to thank Gartner for the recognition.</p><h2 id="cloudflare-waap">Cloudflare WAAP</h2><p>At Cloudflare, we have built several features that fall under the Web Application and API Protection (WAAP) umbrella.</p><h3 id="ddos-protection-mitigation">DDoS protection &amp; mitigation</h3><p>Our <a href="https://www.cloudflare.com/network/">network</a>, which spans more than 275 cities in over 100 countries is the backbone of our platform, and is a core component that allows us to mitigate <a href="http://blog.cloudflare.com/ddos-attack-trends-for-2022-q2/">DDoS attacks of any size</a>.</p><p>To help with this, our network is intentionally anycasted and advertises the same IP addresses from all locations, allowing us to “split” incoming traffic into manageable chunks that each location can handle with ease, and this is especially important when mitigating large volumetric Distributed Denial of Service (DDoS) attacks.</p><p>The system is designed to require little to no configuration while also being “always-on” ensuring attacks are mitigated instantly. Add to that some very smart software such as our new <a href="http://blog.cloudflare.com/location-aware-ddos-protection/">location aware mitigation</a>, and DDoS attacks become a solved problem.</p><p>For customers with very specific traffic patterns, <a href="http://blog.cloudflare.com/http-ddos-managed-rules/">full configurability of our DDoS Managed Rules</a> is just a click away.</p><h3 id="web-application-firewall">Web Application Firewall</h3><p>Our <a href="https://www.cloudflare.com/waf/">WAF</a> is a core component of our application security and ensures hackers and vulnerability scanners have a hard time trying to find potential vulnerabilities in web applications.</p><p>This is very important when zero-day vulnerabilities become publicly available as we’ve seen bad actors attempt to leverage new vectors within hours of them becoming public. <a href="http://blog.cloudflare.com/tag/log4j/">Log4J</a>, and even more recently the <a href="http://blog.cloudflare.com/cloudflare-customers-are-protected-from-the-atlassian-confluence-cve-2022-26134/">Confluence CVE</a>, are just two examples where we observed this behavior. That’s why our WAF is also backed by a team of security experts who <a href="https://developers.cloudflare.com/waf/change-log/scheduled-changes">constantly monitor and develop/improve signatures</a> to ensure we “buy” precious time for our customers to harden and patch their backend systems when necessary. Additionally, and complementary to signatures, our <a href="http://blog.cloudflare.com/waf-ml/">WAF machine learning system</a> classifies each request providing a much wider view in traffic patterns.</p><p>Our WAF comes packed with many advanced features such as <a href="https://developers.cloudflare.com/waf/exposed-credentials-check/">leaked credential checks</a>, <a href="https://developers.cloudflare.com/waf/analytics/">advanced analytics</a> and <a href="http://blog.cloudflare.com/get-notified-when-your-site-is-under-attack/">alerting</a> and <a href="https://developers.cloudflare.com/waf/managed-rulesets/payload-logging/">payload logging</a>.</p><h3 id="bot-management">Bot Management</h3><p>It is no secret that <a href="https://radar.cloudflare.com/">a large portion of web traffic is automated</a>, and while not all automation is bad, some is unnecessary and may also be malicious.</p><p>Our <a href="https://www.cloudflare.com/products/bot-management/">Bot Management</a> product works in parallel to our WAF and scores every request with the likelihood of it being generated by a bot, allowing you to easily filter unwanted traffic by deploying a WAF Custom Rule, all this backed by powerful analytics. We make this easy by also maintaining a list of <a href="https://radar.cloudflare.com/verified-bots">verified bots</a> that can be used to further improve a security policy.</p><p>In the event you want to block automated traffic, <a href="http://blog.cloudflare.com/end-cloudflare-captcha/">Cloudflare's managed challenge</a> ensures that only bots receive a hard time without impacting the experience of real users.</p><h3 id="api-gateway">API Gateway</h3><p>API traffic, by definition, is very well-structured relative to standard web pages consumed by browsers. At the same time, APIs tend to be closer abstractions to back end databases and services, resulting in increased attention from malicious actors and often go unnoticed even to internal security teams (shadow APIs).</p><p><a href="http://blog.cloudflare.com/api-gateway/">API Gateway</a>, that can be layered on top of our WAF, helps you both <a href="https://developers.cloudflare.com/api-shield/security/api-discovery/">discover API endpoints</a> served by your infrastructure, as well detect potential anomalies in traffic flows that may indicate compromise, both from a <a href="https://developers.cloudflare.com/api-shield/security/volumetric-abuse-detection/">volumetric</a> and <a href="https://developers.cloudflare.com/api-shield/security/sequential-abuse-detection/">sequential</a> perspective.</p><p>The nature of APIs also allows API Gateway to much more easily provide a positive security model contrary to our WAF: only allow known good traffic and block everything else. Customers can leverage <a href="https://developers.cloudflare.com/api-shield/security/schema-validation/">schema protection</a> and <a href="https://developers.cloudflare.com/api-shield/security/mtls/">mutual TLS authentication (mTLS)</a> to achieve this with ease.</p><h3 id="page-shield">Page Shield</h3><p>Attacks that leverage the browser environment directly can go unnoticed for some time, as they don’t necessarily require the back end application to be compromised. For example, if any third party JavaScript library used by a web application is performing malicious behavior, application administrators and users may be none the wiser while credit card details are being leaked to a third party endpoint controlled by an attacker. This is a common vector for Magecart, one of many client side security attacks.</p><p><a href="https://www.cloudflare.com/page-shield/">Page Shield</a> is solving client side security by providing active monitoring of third party libraries and <a href="https://developers.cloudflare.com/page-shield/reference/alerts/">alerting application owners whenever a third party asset shows malicious activity</a>. It leverages both public standards such as content security policies (CSP) along with <a href="http://blog.cloudflare.com/detecting-magecart-style-attacks-for-pageshield/">custom classifiers</a> to ensure coverage.</p><p>Page Shield, just like our other WAAP products, is fully integrated on the Cloudflare platform and requires one single click to turn on.</p><h3 id="security-center">Security Center</h3><p>Cloudflare's new <a href="https://www.cloudflare.com/securitycenter/">Security Center</a> is the home of the WAAP portfolio. A single place for security professionals to get a broad view across both <a href="https://developers.cloudflare.com/security-center/tasks/review-insights/">network</a> and <a href="https://developers.cloudflare.com/security-center/tasks/review-infrastructure/">infrastructure</a> assets protected by Cloudflare.</p><p>Moving forward we plan for the Security Center to be the starting point for forensics and analysis, allowing you to also leverage Cloudflare threat intelligence <a href="http://blog.cloudflare.com/security-center-investigate/">when investigating incidents</a>.</p><h2 id="the-cloudflare-advantage">The Cloudflare advantage</h2><p>Our WAAP portfolio is delivered from a single horizontal platform, allowing you to leverage all security features without additional deployments. Additionally, scaling, maintenance and updates are fully managed by Cloudflare allowing you to focus on delivering business value on your application.</p><p>This applies even beyond WAAP, as, although we started building products and services for web applications, our position in the network allows us to protect anything connected to the Internet, including teams, offices and internal facing applications. All from the same single platform. Our <a href="https://www.cloudflare.com/products/zero-trust/">Zero Trust portfolio</a> is now an integral part of our business and WAAP customers can start leveraging our secure access service edge (SASE) with just a few clicks.</p><p>If you are looking to consolidate your security posture, both from a management and budget perspective, application services teams can use the same platform that internal IT services teams use, to protect staff and internal networks.</p><h2 id="continuous-innovation">Continuous innovation</h2><p>We did not build our WAAP portfolio overnight, and over just the past year we’ve released more than five major WAAP portfolio security product releases. To showcase our speed of innovation, here is a selection of our top picks:</p><ul><li><a href="http://blog.cloudflare.com/protecting-apis-from-abuse-and-data-exfiltration/">API Shield Schema Protection</a>: traditional signature based WAF approaches (negative security model) don’t always work well with well-structured data such as API traffic. Given the fast growth in API traffic across the network we built a new incremental product that allows you to enforce API schemas directly at the edge using a positive security model: only let well-formed data through to your origin web servers;</li><li><a href="http://blog.cloudflare.com/api-abuse-detection/">API Abuse Detection</a>: complementary to API Schema Protection, API Abuse Detection warns you whenever anomalies are detected on your API endpoints. These can be triggered by unusual traffic flows or patterns that don’t follow normal traffic activity;</li><li><a href="http://blog.cloudflare.com/new-cloudflare-waf/">Our new Web Application Firewall</a>: built on top of our new Edge Rules Engine, the core Web Application Firewall received a complete overhaul, all the way from engine internals to the UI. Better performance both in terms of latency and efficacy at blocking malicious payloads, along with brand-new capabilities including but not limited to Exposed Credential Checks, account wide configurations and payload logging;</li><li><a href="http://blog.cloudflare.com/http-ddos-managed-rules/">DDoS customizable Managed Rules</a>: to provide additional configuration flexibility, we started exposing some of our internal DDoS mitigation managed rules for custom configurations to further reduce false positives and allow customers to increase thresholds / detections as required;</li><li><a href="http://blog.cloudflare.com/security-center/">Security Center</a>: Cloudflare view on infrastructure and network assets, along with alerts and notifications for miss configurations and potential security issues;</li><li><a href="http://blog.cloudflare.com/page-shield-generally-available/">Page Shield</a>: based on growing customer demand and the rise of attack vectors focusing on the end user browser environment, Page Shield helps you detect whenever malicious JavaScript may have made its way into your application’s code;</li><li><a href="http://blog.cloudflare.com/api-gateway/">API Gateway</a>: full API management, including routing directly from the Cloudflare edge, with API Security baked in, including encryption and mutual TLS authentication (mTLS);</li><li><a href="http://blog.cloudflare.com/waf-ml/">Machine Learning WAF</a>: complementary to our WAF Managed Rulesets, our new ML WAF engine, scores every single request from 1 (clean) to 99 (malicious) giving you additional visibility in both valid and non-valid malicious payloads increasing our ability to detect targeted attacks and scans towards your application;</li></ul><h2 id="looking-forward">Looking forward</h2><p>Our roadmap is packed with both new application security features and improvements to existing systems. As we learn more about the Internet we find ourselves better equipped to keep your applications safe. Stay tuned for more.</p><!--kg-card-begin: markdown--><p><small><em>Gartner, “Magic Quadrant for Web Application and API Protection”, Analyst(s): Jeremy D'Hoinne, Rajpreet Kaur, John Watts, Adam Hils, August 30, 2022.</em></small></p>
<p><small><em>Gartner and Magic Quadrant are registered trademarks of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved.<br>
Gartner does not endorse any vendor, product or service depicted in its research publications, and does not advise technology users to select only those vendors with the highest ratings or other designation.</em></small></p>
<p><small><em>Gartner research publications consist of the opinions of Gartner’s research organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this research, including any warranties of merchantability or fitness for a particular purpose.</em></small></p>
<!--kg-card-end: markdown-->]]></content:encoded></item><item><title><![CDATA[Improving the accuracy of our machine learning WAF using data augmentation and sampling]]></title><description><![CDATA[Data Generation and Augmentation methods to train an effective machine learning WAF]]></description><link>https://blog.cloudflare.com/data-generation-and-sampling-strategies/</link><guid isPermaLink="false">63125f065f22d3000b330518</guid><category><![CDATA[WAF]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[AI]]></category><category><![CDATA[Edge Inference]]></category><category><![CDATA[Data Science]]></category><dc:creator><![CDATA[Vikram Grover]]></dc:creator><pubDate>Mon, 05 Sep 2022 13:00:00 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/09/image2-1.png" medium="image"/><content:encoded><![CDATA[<figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/image2-3.png" class="kg-image" alt="Improving the accuracy of our machine learning WAF using data augmentation and sampling"></figure><img src="http://blog.cloudflare.com/content/images/2022/09/image2-1.png" alt="Improving the accuracy of our machine learning WAF using data augmentation and sampling"><p>At Cloudflare, we are always looking for ways to make our customers' faster and more secure. A key part of that commitment is our ongoing investment in research and development of new technologies, such as the work on our machine learning based <a href="https://www.cloudflare.com/en-gb/learning/ddos/glossary/web-application-firewall-waf/">Web Application Firewall (WAF)</a> solution we announced during <a href="http://blog.cloudflare.com/waf-ml/">security week</a>.</p><p>In this blog, we’ll be discussing some of the data challenges we encountered during the machine learning development process, and how we addressed them with a combination of data augmentation and generation techniques.</p><p>Let’s jump right in!</p><h2 id="introduction">Introduction</h2><p>The purpose of a WAF is to analyze the characteristics of a HTTP request and determine whether the request contains any data which may cause damage to destination server systems, or was generated by an entity with malicious intent. A WAF typically protects applications from common attack vectors such as <a href="https://www.cloudflare.com/learning/security/threats/cross-site-scripting/">cross-site-scripting (XSS)</a>, file inclusion and <a href="https://www.cloudflare.com/learning/security/threats/sql-injection/">SQL injection</a>, to name a few. These attacks can result in the loss of sensitive user data and damage to critical software infrastructure, leading to monetary loss and reputation risk, along with direct harm to customers.</p><h3 id="how-do-we-use-machine-learning-for-the-waf">How do we use machine learning for the WAF?</h3><p>The Cloudflare ML solution, at a high level, trains a classifier to distinguish between various traffic types and attack vectors, such as SQLi, XSS, Command Injection, etc. based on structural or statistical properties of the content. This is achieved by performing the following operations:</p><ol><li>We inspect the raw HTTP input and perform some number of transformations on it such as normalization, content substitutions, or de-duplication.</li><li>Decompose or partition it via some process of <a href="https://en.wikipedia.org/wiki/Lexical_analysis#Tokenization">tokenization</a>, generate statistical information about the content, or extract structural data.</li><li>Compute optimal internal numerical representations of the inputs via the process of training the model. The nature of these internal representations depends on the class of model and architecture.</li><li>Learn to map internal content representations against <a href="https://developers.google.com/machine-learning/glossary#class">classes</a> (XSS, SQLi or others), scores or some other target of interest.</li><li>At run-time, use previously learned representations and mappings to analyze a new input and provide the most likely label or score for it. The score ranges from 1 to 99, with 1 indicating that the request is almost certainly malicious and 99 indicating that the request is probably clean.</li></ol><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/image3-1.png" class="kg-image" alt="Improving the accuracy of our machine learning WAF using data augmentation and sampling"></figure><p>This reasonable starting point stumbles immediately upon a critical challenge right from the start: we need high quality labeled data, and lots of it as that has the biggest impact on model performance. Contrary to well-researched fields like image recognition, text sentiment analysis, or classification, large datasets of HTTP requests with malicious payloads embedded are difficult to get. </p><p>To make matters even harder, strict implementation requirements for a production-quality WAF restrict the complexity of our potential ML models or architectures to ones that are relatively simple and light-weight, implying that we cannot simply pave over shortcomings of the data.</p><h2 id="data-and-challenges">Data and challenges</h2><p>The selection of a dataset is likely the most difficult of all the aspects that contribute to the final set of attributes of a machine learning model. In most cases, the model is tasked with learning the distribution of the data in some statistical sense, thus choosing and curating the dataset to ensure that the desired properties of the final solution are even possible to learn is incredibly crucial! ML models are only as reliable as the data used to train them. If we train an ML model on an incomplete dataset, or on data that doesn’t accurately represent the population, predictions might be inaccurate as they will be a direct reflection of the data. </p><p>To build a strong ML WAF, a good dataset must have large volumes of heterogeneous data covering malicious samples for all attack categories, a diverse set of negative/benign samples, and samples representing a broad spectrum of obfuscation techniques.</p><p>Due to those constraints, creating a solid dataset has a number of challenges:</p><h3 id="privacy">Privacy</h3><p>Privacy requirements limit data availability and how it can be used. Cloudflare has strict privacy guidelines and does not keep all request data - it simply isn't available, and what is available must be carefully selected, anonymised, and stripped of sensitive information. </p><h3 id="heterogeneity-of-samples">Heterogeneity of samples</h3><p>Due to the wide assortment of potential request content types and forms, finding enough benign samples is difficult. Furthermore, it is challenging to collect data that represents requests with various charsets and content-encodings. Covering all attack configurations is also important because some attacks can be inserted into essentially any kind of request (e.g. five bytes in a huge "regular" request)</p><h3 id="sample-difficulty">Sample difficulty</h3><p>We want a dataset with a good mix of attack techniques and isn’t dominated by the ones that are easily generated by tools which simply swap out constants, transform expressions through invariants, and so on (sqli-fuzzer). Additionally, the vast majority of freely available samples in the wild are fairly trivial auto-generated payloads as part of indiscriminate scanning and discovery tools. They have very similar structural and statistical characteristics. Some of them are fairly old as well and do not reflect the current software landscape. How to "grade" the sample difficulty is not immediately obvious! What’s easy to a human may not be easy for a particular preprocessor/model, and vice-versa.</p><h3 id="noisy-labels">Noisy labels</h3><p>Label noise affects results a lot, especially when it comes to esoteric, specific, or unusual attacks which are likely to be classified as benign by rules WAF.</p><p>What’s the strategy to overcome this?</p><h2 id="data-augmentation">Data augmentation</h2><p>In simple terms, <a href="https://en.wikipedia.org/wiki/Data_augmentation">Data Augmentation</a> is a process of generating artificial (but realistic) data to increase the diversity of our data by studying statistical distribution of existing real-world data.</p><p>This is crucial for us because one of the biggest concerns with rules-based WAFs is <a href="https://en.wikipedia.org/wiki/False_positives_and_false_negatives">false positives</a>. False positives are a serious challenge for WAFs because the risk of accidentally filtering legitimate traffic deters users from employing very strict rulesets. Data augmentation is used to build a solution that does not rely on observing specific high-risk keywords or character sequences, but instead uses a more holistic analysis of content and context, making it considerably less likely to block legitimate requests. </p><p>There are many sequences of characters which appear almost exclusively in payloads, but are themselves not dangerous. In order to reduce false positives and improve overall performance, we focussed on generating a lot of heterogeneous negative samples to force the model to consider the structural, semantic, and statistical properties of the content when making a classification decision.</p><p>In the context of our data and use cases, data augmentation means that we mutate benign content in a variety of ways as the content will remain benign (this isn’t going to accidentally turn it into a valid payload, with probability 1). For instance, we can add random character noise, permute keywords, merge benign content together from multiple sources, and so on. Alternatively, we can seed benign content with ‘dangerous’ keywords or <a href="https://en.wikipedia.org/wiki/N-gram">ngrams</a> frequently occuring in payloads - this results in a benign sample, but ideally will teach the model not to be too sensitive to the presence of malicious tokens lacking the proper semantics and structure.</p><h3 id="benign-content">Benign content</h3><p>First and foremost, generating benign content is way easier. Mutating a malicious block of content into different malicious blocks is difficult because malicious payloads have a stricter grammar and syntax than general HTTP content due to the fact that it has code, therefore they must be manipulated in a specific manner. </p><p>However, there are a few options  if we want to do this in the future. Tools like sqli-fuzzer,  automates the process of fuzzing a given payload by applying transformations which preserve the underlying semantics while changing the representation or adding obfuscation. Outside existing third-party tools, it's possible to generate our own malicious payloads using various "append malicious content to non-malicious content" techniques, with the trade off that this doesn't actually generate *new* malicious content, just puts it into a different context.</p><h3 id="pseudo-random-noise-samples">Pseudo-random noise samples</h3><p>A useful approach we identified for bolstering the number of negative training samples was to generate large quantities of pseudo-random strings of increasing complexity.</p><p>The probability of any pseudo-random string (drawn from essentially any token distribution) being a valid payload or malicious attack is essentially zero, but we can build a series of token sampling distributions that make it increasingly difficult for the model to distinguish them from a real payload, and we discovered that this resulted in dramatically better performance in terms of <a href="https://en.wikipedia.org/wiki/False_positive_rate">false positive rate</a>, robustness, and overall model properties.</p><p>This approach works by taking a collection of tokens and a probability distribution over these tokens, and independently sampling a stream of tokens from it to create our ‘sample’. Each sample length is selected from a separate discrete sample length distribution.</p><p>For an extremely simple example, we could take a token collection consisting of ASCII characters and a uniform sampling distribution:</p><p><code>['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i', 'j', 'k', 'l', 'm', 'n', 'o', 'p', 'q', 'r', 's', 't', 'u', 'v', 'w', 'x', 'y', 'z', '0', '1', '2', '3', '4', '5', '6', '7', '8', '9']</code></p><p>We sample random strings of length 0-32 from this to get some (uninteresting) negative samples:</p><p><em><code>8hwk1d740hfstbb4aogbpi4qayppvdl41b6blornuzktp4yl</code></em></p><p><em><code>1deq7rug1zftmn9tjr73yttjnye99zh2140z2x9lr8n6sxhucdgn6bmqvfv7auw8fwbkrtxilk45ht-</code></em></p><p>We wouldn’t expect even a very simple model to struggle to learn that these samples are benign,  but as we increase the complexity of the token collections, we can move towards much more ‘difficult’ noise examples, including elements such as: fragments of valid URIs, user agents, XML/XSLT content or even restricted language identifiers, or keywords.</p><p>Here are some examples of more complex token collections and the kinds of random strings they produce as our negative samples:</p><p><em>Ascii_script: alphanumeric characters plus  '&lt;', '&gt;', '/', '&lt;/', '-', '+', '=', '&lt; ', ' &gt;', ' ', ' /&gt;'</em></p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screenshot-2022-09-05-at-13.16.14.png" class="kg-image" alt="Improving the accuracy of our machine learning WAF using data augmentation and sampling"></figure><p><em>alphanumerics, plus special characters, plus a variant of full javascript or sql keywords and (multi-character) sub-token fragments</em></p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screenshot-2022-09-05-at-13.17.48.png" class="kg-image" alt="Improving the accuracy of our machine learning WAF using data augmentation and sampling"></figure><p>It’s fairly straightforward to construct a suite of these noise generators of varying complexity, and targeting different types of content: JSON, XML, URIs with SQL-esque ‘noise’, and so on. As the strings get sufficiently long, the probability that they will contain at least some dangerous looking subsequences grows, so it’s also an excellent test of model robustness. </p><p>We make extensive use of noise strings to enhance the core dataset used for training and testing the model by directly training the model on increasingly difficult noise before fine-tuning on exclusively real data, appending noise of varying complexity to malicious(real) samples or benign samples to both induce and test for model robustness for padding attacks, and estimating false positive rate for certain classes of benign content.</p><h3 id="beyond-independent-sampling-of-random-strings">Beyond independent sampling of random strings?</h3><p>A natural extension to the above method for generating pseudo-random strings is to drop the ‘independence’ assumption for sampling tokens. This means that we’re starting to emulate the process by which real data is generated, to some extent, yielding samples with increasingly realistic local (and eventually global) structure. Some approaches for this might include a simple Markov chain, and extend all the way to state-of-the-art Large Language Models.</p><p>We experimented with using contemporary autoregressive language models trained on our corpus of real malicious payloads and found it extremely effective at generating novel payloads, as well as transforming payloads into sophisticated obfuscated representations. As the language models approached convergence on the data the likelihood of each sample being a valid payload approached 100%, allowing us to use early samples as ‘extremely strong negatives’ and the later samples as positive samples. The success of this work has suggested that deeper investigation into the use of language models for security analysis may be fruitful, not only for training classifiers, but also for creating powerful adversarial pen-testing agents.</p><h2 id="results-summary">Results summary</h2><p>Let’s see a comparative summary of results and improvements, before and after the augmentation:</p><h3 id="model-performance-on-evaluation-metrics">Model performance on evaluation metrics</h3><p>The effectiveness of machine learning models for classification problems can be evaluated using a wide range of metrics, including <a href="https://en.wikipedia.org/wiki/Evaluation_of_binary_classifiers">accuracy, precision, recall, F1 Score,</a> and others. It is important to note that in addition to using quantitative metrics, we also consider the model's general properties and behavioral constraints. This criteria and metrics-based approach is especially important in our domain where data is inherently noisy, labels are not trustworthy, the domain of the inputs is extremely large, and hard to cover with samples. </p><p>For this post, we will concentrate on key quantitative metrics like F1 score even though we examine a variety of metrics to assess the model performance. F1 score is the weighted average (harmonic mean) of precision and recall. We can represent the F1 score with the formula:</p><figure class="kg-card kg-image-card"><img src="https://lh4.googleusercontent.com/57sg9aMxBhrWv1CnYdMvRNodVVjcvCUZIrbRExH2sD4-f-7sy2mOKjqRaU3Z6uAH_WEVXHLmzuHBooN_ThqEfan7dVu5khiwLzr8a7ts1UpyNdP6bVWQ4WNhxv_o98jY0642lFynuCCQ1_WiWOMoPy4BVgI8jsCW70gDwvfzkxkTBTpgjHk8Or47lw" class="kg-image" alt="Improving the accuracy of our machine learning WAF using data augmentation and sampling"></figure><p>Where,</p><p><strong>True Positives (TP):</strong> malicious content classified correctly by the model</p><p><strong>False Positives (FP):</strong> benign content that the model classified as malicious</p><p><strong>True Negatives (TN):</strong> benign content classified correctly by the model</p><p><strong>False Negatives (FN):</strong> malicious content that the model classified as benign</p><p>Since this formula takes false positives and false negatives into consideration, this score is more reliable than other metrics. There are a few methods to calculate this for <a href="https://en.wikipedia.org/wiki/Multiclass_classification">multi-class</a> problems, like <a href="https://towardsdatascience.com/multi-class-metrics-made-simple-part-ii-the-f1-score-ebe8b2c2ca1">Macro F1 Score, Micro F1 Score and Weighted F1 Score</a>. Although each method has advantages and disadvantages, we obtained nearly identical results with all three methods. Below are the numbers:</p><!--kg-card-begin: html--><style type="text/css">
.tg  {border-collapse:collapse;border-spacing:0;}
.tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;
  overflow:hidden;padding:10px 5px;word-break:normal;}
.tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;
  font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;}
.tg .tg-5t9h{background-color:#A4C2F4;font-weight:bold;text-align:center;vertical-align:top}
.tg .tg-m0cw{background-color:#C9DAF8;font-weight:bold;text-align:center;vertical-align:top}
.tg .tg-p7vi{background-color:#C9DAF8;font-weight:bold;text-align:left;vertical-align:top}
.tg .tg-skn0{background-color:#A4C2F4;text-align:left;vertical-align:top}
.tg .tg-ktyi{background-color:#FFF;text-align:left;vertical-align:top}
.tg .tg-7yig{background-color:#FFF;text-align:center;vertical-align:top}
</style>
<table class="tg" width="100%">
<thead>
  <tr>
    <th class="tg-skn0" rowspan="2"></th>
    <th class="tg-5t9h" colspan="3" rowspan="2"><span style="font-weight:700;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Without Augmentation</span></th>
    <th class="tg-5t9h" colspan="3" rowspan="2"><span style="font-weight:700;font-style:normal;text-decoration:none;color:#000;background-color:transparent">With Augmentation</span></th>
  </tr>
  <tr>
  </tr>
</thead>
<tbody>
  <tr>
    <td class="tg-p7vi"><span style="font-weight:700;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Class</span></td>
    <td class="tg-m0cw"><span style="font-weight:700;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Precision</span></td>
    <td class="tg-m0cw"><span style="font-weight:700;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Recall</span></td>
    <td class="tg-m0cw"><span style="font-weight:700;font-style:normal;text-decoration:none;color:#000;background-color:transparent">F1 Score</span></td>
    <td class="tg-m0cw"><span style="font-weight:700;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Precision</span></td>
    <td class="tg-m0cw"><span style="font-weight:700;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Recall</span></td>
    <td class="tg-m0cw"><span style="font-weight:700;font-style:normal;text-decoration:none;color:#000;background-color:transparent">F1 Score</span></td>
  </tr>
  <tr>
    <td class="tg-ktyi"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Benign</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.69</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.17</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.27</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.98</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">1.00</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.99</span></td>
  </tr>
  <tr>
    <td class="tg-ktyi"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">SQLi</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.77</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.96</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.85</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">1.00</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">1.00</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">1.00</span></td>
  </tr>
  <tr>
    <td class="tg-ktyi"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">XSS</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.56</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.94</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.70</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">1.00</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.98</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.99</span></td>
  </tr>
  <tr>
    <td class="tg-ktyi"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Total(Micro Average)</span></td>
    <td class="tg-7yig"></td>
    <td class="tg-7yig"></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.67</span></td>
    <td class="tg-7yig"></td>
    <td class="tg-7yig"></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.99</span></td>
  </tr>
  <tr>
    <td class="tg-ktyi"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Total(Macro Average)</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.67</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.69</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.61</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.99</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.99</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.99</span></td>
  </tr>
  <tr>
    <td class="tg-ktyi"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Total(Weighted Average)</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.68</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.67</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.60</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.99</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.99</span></td>
    <td class="tg-7yig"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">0.99</span></td>
  </tr>
</tbody>
</table><!--kg-card-end: html--><p>The important takeaway is that the range of this F1 score is best at 1 and worst at 0.</p><p>The model after augmentation appears to have similar precision and recall with good overall performance, as indicated by a value of 0.99 after augmentation, compared to 0.61 for Macro F1.</p><p>So far in the results summary, we've only discussed F1 Score; however, there are other improvements in characteristics that we've observed in the model that are listed below:</p><p><strong>False positive characteristics</strong></p><ul><li>Estimated false positive rate reduced by approximately 80% on test data sets. There are significantly fewer false positives involving PromQL and other SQL-structured analogues. PromQL examples result in high scores and are classified correctly:</li></ul><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screenshot-2022-09-05-at-12.35.06.png" class="kg-image" alt="Improving the accuracy of our machine learning WAF using data augmentation and sampling"></figure><p>Today, the only major category of false positives are literal SQL or JavaScript files.</p><ul><li>General false positive rate on noise from JSON-esque, XML/SOAP-esque, and SQL-esque content-generators reduced to about a 1/100,000 rate from about 1/50 to 1/1.</li></ul><p><strong>True positive characteristics</strong></p><ul><li><strong>True positive</strong> rate for highly fuzzed content is vastly improved. Models trained solely on real data were easily bypassed by advanced fuzzing tools, whereas models trained on real plus augmented data are extremely resistant, with many payloads receiving higher risk scores as fuzzing increases. Examples:</li></ul><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screenshot-2022-09-05-at-12.36.25.png" class="kg-image" alt="Improving the accuracy of our machine learning WAF using data augmentation and sampling"></figure><p>These yield approximately same scores as they are a result of only a few byte   alterations</p><ul><li>Proportion of client-provided test sets that primarily contain payloads not blocked by rules-waf for XSS/SQLi successfully classified is about 97.5% (with the remaining 2.5% being arguable) up from about 91%.<br></li><li>Padding a payload with almost any amount of ASCII, JSON-esque, special-characters, or other content will not reduce the risk score substantially. Due to the addition of hard noise long length augmented training samples, even a six byte payload in a 100 kilobyte string will be caught. Examples:</li></ul><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screenshot-2022-09-05-at-12.37.03.png" class="kg-image" alt="Improving the accuracy of our machine learning WAF using data augmentation and sampling"></figure><p>They both generate similar scores even though the latter has junk padding around the payload.</p><p><strong>Execution performance</strong></p><ul><li>Runtime characteristics are unchanged for inference.</li></ul><p>On top of that, we validated the model against the Cloudflare’s highly mature signature-based WAF and confirmed that machine learning WAF performs comparable to signature WAF, with the ML WAF demonstrating its strength particularly in cases of correctly handling highly obfuscated or irregularly fuzzed content (as well as avoiding some rules-based engine false positives). ​​Finally, we were able to conclude that augmentation helps in improving the model performance and induce the right set of properties.</p><h2 id="conclusion">Conclusion</h2><p>We built a machine learning powered WAF, with the substantial challenge to gather a diversified training set, given constraints to avoid sensitive real customer data for privacy and regulatory considerations. To create a broader and diversified dataset without requiring vast amounts of sensitive data, we used techniques such as fuzzing, data augmentation, and synthetic data generation. This allowed us to improve the solution's false positive robustness and overall model performance.</p><p>Furthermore, these techniques reduced the time complexity required to retrieve/clean real data, and helped induce the correct model behavior. In the future, we intend to investigate autoregressive language models to generate synthetic pseudo-valid payloads.</p>]]></content:encoded></item><item><title><![CDATA[Blocking Kiwifarms]]></title><description><![CDATA[We have blocked Kiwifarms. Visitors to any of the Kiwifarms sites that use any of Cloudflare's services will see a Cloudflare block page and a link to this post. ]]></description><link>https://blog.cloudflare.com/kiwifarms-blocked/</link><guid isPermaLink="false">6313ba71ec6c4e000bf2af8d</guid><category><![CDATA[Abuse]]></category><category><![CDATA[Legal]]></category><dc:creator><![CDATA[Matthew Prince]]></dc:creator><pubDate>Sat, 03 Sep 2022 22:15:35 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/09/The-Cloudflare-Blog-1.png" medium="image"/><content:encoded><![CDATA[<!--kg-card-begin: markdown--><img src="http://blog.cloudflare.com/content/images/2022/09/The-Cloudflare-Blog-1.png" alt="Blocking Kiwifarms"><p><em><small>This post is also available in <a href="http://blog.cloudflare.com/zh-cn/kiwifarms-blocked-zh-cn/">简体中文</a>, <a href="http://blog.cloudflare.com/ja-jp/kiwifarms-blocked-ja-jp/">日本語</a>, <a href="http://blog.cloudflare.com/de-de/kiwifarms-blocked-de-de/">Deutsch</a>, <a href="http://blog.cloudflare.com/fr-fr/kiwifarms-blocked-fr-fr/">Français</a>, <a href="http://blog.cloudflare.com/es-es/kiwifarms-blocked-es-es/">Español</a>.</small></em></p>
<!--kg-card-end: markdown--><p>We have blocked Kiwifarms. Visitors to any of the Kiwifarms sites that use any of Cloudflare's services will see a Cloudflare block page and a link to this post. Kiwifarms may move their sites to other providers and, in doing so, come back online, but we have taken steps to block their content from being accessed through our infrastructure.</p><p>This is an extraordinary decision for us to make and, given Cloudflare's role as an Internet infrastructure provider, a dangerous one that we are not comfortable with. However, the rhetoric on the Kiwifarms site and specific, targeted threats have escalated over the last 48 hours to the point that we believe there is an unprecedented emergency and immediate threat to human life unlike we have previously seen from Kiwifarms or any other customer before.</p><h3 id="escalating-threats">Escalating threats</h3><p>Kiwifarms has frequently been host to revolting content. Revolting content alone does not create an emergency situation that necessitates the action we are taking today. Beginning approximately two weeks ago, a pressure campaign started with the goal to deplatform Kiwifarms. That pressure campaign targeted Cloudflare as well as other providers utilized by the site.</p><p>Cloudflare provided security services to Kiwifarms, protecting them from DDoS and other cyberattacks. We have never been their hosting provider. <a href="http://blog.cloudflare.com/cloudflares-abuse-policies-and-approach/">As we outlined last Wednesday</a>, we do not believe that terminating security services is appropriate, even to revolting content. In a law-respecting world, the answer to even illegal content is not to use other illegal means like DDoS attacks to silence it.</p><p>We are also not taking this action directly because of the pressure campaign. While we have empathy for its organizers, we are committed as a security provider to protecting our customers even when they run deeply afoul of popular opinion or even our own morals. The <a href="http://blog.cloudflare.com/cloudflares-abuse-policies-and-approach/">policy we articulated last Wednesday remains our policy</a>. We continue to believe that the best way to relegate cyberattacks to the dustbin of history is to give everyone the tools to prevent them.</p><p>However, as the pressure campaign escalated, so did the rhetoric on the Kiwifarms site. Feeling attacked, users of Kiwifarms became even more aggressive. Over the last two weeks, we have proactively reached out to law enforcement in multiple jurisdictions highlighting what we believe are potential criminal acts and imminent threats to human life that were posted to the site.</p><h3 id="legal-process">Legal process</h3><p>While law enforcement in these areas are working to investigate what we and others reported, unfortunately the process is moving more slowly than the escalating risk. While we believe that in every other situation we have faced — including the Daily Stormer and 8chan — it would have been appropriate as an infrastructure provider for us to wait for legal process, in this case the imminent and emergency threat to human life which continues to escalate causes us to take this action.</p><p>Hard cases make bad law. This is a hard case and we would caution anyone from seeing it as setting precedent. The <a href="http://blog.cloudflare.com/cloudflares-abuse-policies-and-approach/">policies we articulated last Wednesday remain our policies</a>. For an infrastructure provider like Cloudflare, legal process is still the correct way to deal with revolting and potentially illegal content online.</p><p>But we need a mechanism when there is an emergency threat to human life for infrastructure providers to work expediently with legal authorities in order to ensure the decisions we make are grounded in due process. Unfortunately, that mechanism does not exist and so we are making this uncomfortable emergency decision alone.</p><h3 id="not-the-end">Not the end</h3><p>Finally, we are aware and concerned that our action may only fan the flames of this emergency. Kiwifarms itself will most likely find other infrastructure that allows them to come back online, as the Daily Stormer and 8chan did themselves after we terminated them. And, even if they don't, the individuals that used the site to increasingly terrorize will feel even more isolated and attacked and may lash out further. There is real risk that by taking this action today we may have further heightened the emergency.</p><p>We will continue to work proactively with law enforcement to help with their investigations into the site and the individuals who have posted what may be illegal content to it. And we recognize that while our blocking Kiwifarms temporarily addresses the situation, it by no means solves the underlying problem. That solution will require much more work across society. We are hopeful that our action today will help provoke conversations toward addressing the larger problem. And we stand ready to participate in that conversation.<br></p>]]></content:encoded></item><item><title><![CDATA[Log analytics using ClickHouse]]></title><description><![CDATA[When a request at Cloudflare throws an error, information gets logged in our requests_error pipeline. The error logs are used to help troubleshoot customer-specific or network-wide issues]]></description><link>https://blog.cloudflare.com/log-analytics-using-clickhouse/</link><guid isPermaLink="false">63121b215f22d3000b330462</guid><category><![CDATA[ClickHouse]]></category><category><![CDATA[Deep Dive]]></category><dc:creator><![CDATA[Monika Singh]]></dc:creator><pubDate>Fri, 02 Sep 2022 15:33:24 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-8.30.34-AM.png" medium="image"/><content:encoded><![CDATA[<img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-8.30.34-AM.png" alt="Log analytics using ClickHouse"><p>This is an adapted transcript of a talk we gave at Monitorama 2022. You can find the slides with presenter’s notes <a href="https://docs.google.com/presentation/d/1HLnxh-56LMKZGvrc85qxYUt8-tl4M4izdPbxKj4mdL8/edit?usp=sharing">here</a> and video <a href="https://vimeo.com/730379928?embedded=false&amp;source=vimeo_logo&amp;owner=6548926">here</a>.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.37.21-AM.png" class="kg-image" title="Log analytics using ClickHouse" alt="Log analytics using ClickHouse"></figure><p>When a request at Cloudflare throws an error, information gets logged in our requests_error pipeline. The error logs are used to help troubleshoot customer-specific or network-wide issues.</p><p>We, Site Reliability Engineers (SREs), manage the logging platform. We have been running Elasticsearch clusters for many years and during these years, the log volume has increased drastically. With the log volume increase, we started facing a few issues. Slow query performance and high resource consumption to list a few. We aimed to improve the log consumer's experience by improving query performance and providing cost-effective solutions for storing logs. This blog post discusses challenges with logging pipelines and how we designed the new architecture to make it faster and cost-efficient.</p><p>Before we dive into challenges in maintaining the logging pipelines, let us look at the characteristics of logs.</p><h3 id="characteristics-of-logs">Characteristics of logs</h3><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.37.31-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Unpredictable log volume"></figure><p><strong>Unpredictable</strong> - In today's world, where there are tons of microservices, the amount of logs a centralized logging system will receive is very unpredictable. There are various reasons why capacity estimation of log volume is so difficult. Primarily because new applications get deployed to production continuously, existing applications are automatically scaled up or down to handle business demands or sometimes application owners enable debug log levels and forget to turn it off.</p><p><strong>Semi-structured</strong> - Every application adopts a different logging format. Some are represented in plain-text and others use JSON. The timestamp field within these log lines also varies. Multi-line exceptions and stack traces make them even more unstructured. Such logs add extra resource overhead, requiring additional data parsing and mangling.</p><p><strong>Contextual</strong> - For debugging issues, often contextual information is required, that is, logs before and after an event happened. A single logline hardly helps, generally, it's the group of loglines that helps in building the context. Also, we often need to correlate the logs from multiple applications to draw the full picture. Hence it's essential to preserve the order in which logs get populated at the source.</p><p><strong>Write-heavy</strong> - Any centralized logging system is write-intensive. More than 99% of logs that are written, are never read. They occupy space for some time and eventually get purged by the retention policies. The remaining less than 1% of the logs that are read are very important and we can't afford to miss them.</p><h2 id="logging-pipeline">Logging pipeline</h2><p>Like most other companies, our logging pipeline consists of a producer, shipper, a queue, a consumer and a datastore.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.37.42-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Logging pipeline"></figure><p>Applications (Producers) running on the Cloudflare global network generate the logs. These logs are written locally in Cap'n Proto serialized format. The Shipper (in-house solution) pushes the Cap'n Proto serialized logs through streams for processing to Kafka (queue). We run Logstash (consumer), which consumes from Kafka and writes the logs into ElasticSearch (datastore).The data is then visualized by using Kibana or Grafana. We have multiple dashboards built in both Kibana and Grafana to visualize the data.</p><h2 id="elasticsearch-bottlenecks-at-cloudflare">Elasticsearch bottlenecks at Cloudflare</h2><p>At Cloudflare, we have been running Elasticsearch clusters for many years. Over the years, log volume increased dramatically and while optimizing our Elasticsearch clusters to handle such volume, we found a few limitations.</p><h3 id="mapping-explosion">Mapping Explosion</h3><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.37.50-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Mapping explosion"></figure><p>Mapping Explosion is one of the very well-known limitations of Elasticsearch. Elasticsearch maintains a mapping that decides how a new document and its fields are stored and indexed. When there are too many keys in this mapping, it can take a significant amount of memory resulting in frequent garbage collection. One way to prevent this is to make the schema strict, which means any log line not following this strict schema will end up getting dropped. Another way is to make it semi-strict, which means any field not part of this mapping will not be searchable.</p><h3 id="multi-tenancy-support">Multi-tenancy support</h3><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.38.00-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Bad resource isolation in Elasticsearch"></figure><p>Elasticsearch doesn't have very good multi-tenancy support. One bad user can easily impact cluster performance. There is no way to limit the maximum number of documents or indexes a query can read or the amount of memory an Elasticsearch query can take. A bad query can easily degrade cluster performance and even after the query finishes, it can still leave its impact.</p><h3 id="cluster-operational-tasks">Cluster operational tasks</h3><p>It is not easy to manage Elasticsearch clusters, especially multi-tenant ones. Once a cluster degrades, it takes significant time to get the cluster back to a fully healthy state. In Elasticsearch, updating the index template means reindexing the data, which is quite an overhead. We use hot and cold tiered storage, i.e., recent data in SSD and older data in magnetic drives. While Elasticsearch moves the data from hot to cold storage every day, it affects the read and write performance of the cluster.</p><h3 id="garbage-collection">Garbage collection</h3><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.38.10-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Garbage collection in Elasticsearch"></figure><p>Elasticsearch is developed in Java and runs on a Java Virtual Machine (JVM). It performs garbage collection to reclaim memory that was allocated by the program but is no longer referenced. Elasticsearch requires garbage collection tuning. The default garbage collection in the latest JVM is G1GC. We tried other GC like ZGC, which helped in lowering the GC pause but didn't give us much performance benefit in terms of read and write throughput.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.38.21-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Cloudflare HTTP RPS vs error rate"></figure><p>Elasticsearch is a good tool for full-text search and these limitations are not significant with small clusters, but in Cloudflare, we handle over 35 to 45 million HTTP requests per second, out of which over 500K-800K requests fail per second. These failures can be due to an improper request, origin server errors, misconfigurations by users, network issues and various other reasons.</p><p>Our customer support team uses these error logs as the starting point to triage customer issues. The error logs have a number of fields metadata about various Cloudflare products that HTTP requests have been through. We were storing these error logs in Elasticsearch. We were heavily sampling them since storing everything was taking a few hundreds of terabytes crossing our resource allocation budget. Also, dashboards built over it were quite slow since they required heavy aggregation over various fields. We need to retain these logs for a few weeks per the debugging requirements.</p><h2 id="proposed-solution">Proposed solution</h2><p>We wanted to remove sampling completely, that is, store every log line for the retention period, to provide fast query support over this huge amount of data and to achieve all this without increasing the cost.</p><p>To solve all these problems, we decided to do a proof of concept and see if we could accomplish our requirements using ClickHouse.</p><p>Cloudflare was <a href="http://blog.cloudflare.com/how-cloudflare-analyzes-1m-dns-queries-per-second/">an early adopter of ClickHouse</a> and we have been managing ClickHouse clusters for years. We already had a lot of in-house tooling and libraries for inserting data into ClickHouse, which made it easy for us to do the proof of concept. Let us look at some of the ClickHouse features that make it the perfect fit for storing logs and which enabled us to build our new logging pipeline.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.38.31-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Proposed pipeline"></figure><p>ClickHouse is a column-oriented database which means all data related to a particular column is physically stored next to each other. Such data layout helps in fast sequential scan even on commodity hardware. This enabled us to extract maximum performance out of older generation hardware.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.38.40-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title=" column-oriented database"></figure><p>ClickHouse is designed for analytical workloads where the data has a large number of fields that get represented as ClickHouse columns. We were able to design our new ClickHouse tables with a large number of columns without sacrificing performance.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.38.55-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="ClickHouse designed for analytical workloads"></figure><p>ClickHouse indexes work differently than those in relational databases. In relational databases, the primary indexes are dense and contain one entry per table row. So if you have 1 million rows in the table, the primary index will also have 1 million entries. While In ClickHouse, indexes are sparse, which means there will be only one index entry per a few thousand table rows. ClickHouse indexes enabled us to add new indexes on the fly.</p><p>ClickHouse compresses everything with LZ4 by default. An efficient compression not only helps in minimizing the storage needs but also lets ClickHouse use page cache efficiently.</p><p>One of the cool features of ClickHouse is that the compression codecs can be configured on a per-column basis. We decided to keep default LZ4 compression for all columns. We used special encodings like Double-Delta for the DateTime columns, Gorilla for Float columns and LowCardinality for fixed-size String columns.</p><p>ClickHouse is linearly scalable; that is, the writes can be scaled by adding new shards and the reads can be scaled by adding new replicas. Every node in a ClickHouse cluster is identical. Not having any special nodes helps in scaling the cluster easily.</p><p>Let's look at some optimizations we leveraged to provide faster read/write throughput and better compression on log data.</p><h3 id="inserter">Inserter</h3><p>Having an efficient inserter is as important as having an efficient data store. At Cloudflare, we have been operating quite a few analytics pipelines from where we borrowed most of the concepts while writing our new inserter. We use Cap'n Proto messages as the transport data format since it provides fast data encoding and decoding. Scaling inserters is easy and can be done by adding more Kafka partitions and spawning new inserter pods.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.39.10-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Scaling ClickHouse inserter"></figure><h4 id="batch-size">Batch Size</h4><p>One of the key performance factors while inserting data into ClickHouse is the batch size. When batches are small, ClickHouse creates many small partitions, which it then merges into bigger ones. Thus smaller batch size creates extra work for ClickHouse to do in the background, thereby reducing ClickHouse's performance. Hence it is crucial to set it big enough that ClickHouse can accept the data batch happily without hitting memory limits.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.40.04-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="ClickHouse inserter batch size"></figure><h3 id="data-modeling-in-clickhouse-">Data modeling in ClickHouse.</h3><p>ClickHouse provides in-built sharding and replication without any external dependency. Earlier versions of ClickHouse depended on ZooKeeper for storing replication information, but the recent version removed the ZooKeeper dependency by adding clickhouse-keeper.</p><p>To read data across multiple shards, we use distributed tables, a special kind of table. These tables don't store any data themselves but act as a proxy over multiple underlying tables storing the actual data.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.40.15-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Data modeling in ClickHouse"></figure><p>Like any other database, choosing the right table schema is very important since it will directly impact the performance and storage utilization. We would like to discuss three ways you can store log data into ClickHouse.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.40.32-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Schema modeling in ClickHouse"></figure><p>The first is the simplest and the most strict table schema where you specify every column name and data type. Any logline having a field outside this predefined schema will get dropped. From our experience, this schema will give you the fastest query capabilities. If you already know the list of all possible fields ahead, we would recommend using it. You can always add or remove columns by running ALTER TABLE queries.</p><p>The second schema uses a very new feature of ClickHouse, where it does most of the heavy lifting. You can insert logs as JSON objects and behind the scenes, ClickHouse will understand your log schema and dynamically add new columns with appropriate data type and compression. This schema should only be used if you have good control over the log schema and the number of total fields is less than 1,000. On the one hand it provides flexibility to add new columns as new log fields automatically, but at the same time, one lousy application can easily bring down the ClickHouse cluster.</p><p>The third schema stores all fields of the same data type in one array and then uses ClickHouse inbuilt array functions to query those fields. This schema scales pretty well even when there are more than 1,000 fields, as the number of columns depends on the data types used in the logs. If an array element is accessed frequently, it can be taken out as a dedicated column using the materialized column feature of ClickHouse. We recommend adopting this schema since it provides safeguards against applications logging too many fields.</p><h4 id="data-partitioning">Data partitioning</h4><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.40.47-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Data partitioning"></figure><p>A partition is a unit of ClickHouse data. One common mistake ClickHouse users make is overly granular partitioning keys, resulting in too many partitions. Since our logging pipeline generates TBs of data daily, we created the table partitioned with `toStartOfHour(dateTime).` With this partitioning logic, when a query comes with the timestamp in the WHERE clause, ClickHouse knows the partition and retrieves it quickly. It also helps design efficient data purging rules according to the data retention policies.</p><h4 id="primary-key-selection">Primary key selection</h4><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.40.57-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Primary key selection"></figure><p>ClickHouse stores the data on disk sorted by primary key. Thus, selecting the primary key impacts the query performance and helps in better data compression. Unlike relational databases, ClickHouse doesn't require a unique primary key per row and we can insert multiple rows with identical primary keys. Having multiple primary keys will negatively impact the insertion performance. One of the significant ClickHouse limitations is that once a table is created the primary key can not be updated.</p><h4 id="data-skipping-indexes">Data skipping indexes</h4><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.41.18-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Data skipping indexes"></figure><p>ClickHouse query performance is directly proportional to whether it can use the primary key when evaluating the WHERE clause. We have many columns and all these columns can not be part of the primary key. Thus queries on these columns will have to do a full scan resulting in slower queries. In traditional databases, secondary indexes can be added to handle such situations. In ClickHouse, we can add another class of indexes called data skipping indexes, which uses bloom filters and skip reading significant chunks of data that are guaranteed to have no match.</p><h2 id="abr">ABR</h2><p>We have multiple dashboards built over the requests_error logs. Loading these dashboards were often hitting the memory limits set for the individual query/user in ClickHouse.</p><p>The dashboards built over these logs were mainly used to identify anomalies. To visually identify an anomaly in a metric, the exact numbers are not required, but an approximate number would do. For instance, to understand that errors have increased in a data center, we don’t need the exact number of errors. So we decided to use an in-house library and tool built around a concept called <a href="http://blog.cloudflare.com/explaining-cloudflares-abr-analytics/">ABR</a>.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.41.26-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Adaptive Bitrate"></figure><p>ABR stands for "<a href="https://en.wikipedia.org/wiki/Adaptive_bitrate_streaming">Adaptive Bit Rate</a>" - the term ABR is mainly used in video streaming services where servers select the best resolution for a video stream to match the client and network connection. It is described in great detail in the blog post - <a href="http://blog.cloudflare.com/explaining-cloudflares-abr-analytics/">Explaining Cloudflare's ABR Analytics</a></p><p>In other words, the data is stored at multiple resolutions or sample intervals and the best solution is picked for each query.</p><p>The way ABR works is at the time of writing requests to ClickHouse, it writes the data in a number of tables with different sample intervals. For instance table_1 stores 100% of data, table_10 stores 10% of data, table_100 stores 1% of data and table_1000 stores 0.1% data so on and so forth. The data is duplicated between the tables. Table_10 would be a subset of table_1.</p><h2 id="demo">Demo</h2><p>In Cloudflare, we use in-house libraries and tools to insert data into ClickHouse, but this can be achieved by using an open source tool - vector.dev</p><p>If you would like to test how log ingestion into ClickHouse works, you can refer or use the demo <a href="https://github.com/cloudflare/cloudflare-blog/tree/master/2022-08-log-analytics">here</a>.</p><p>Make sure you have docker installed and run `docker compose up` to get started.</p><p>This would bring up three containers, Vector.dev for generating vector demo logs, writing it into ClickHouse, ClickHouse container to store the logs and Grafana instance to visualize the logs.</p><p>When the containers are up, visit <a href="http://localhost:3000/dashboards">http://localhost:3000/dashboards</a> to play with the prebuilt demo dashboard.</p><h2 id="conclusion">Conclusion</h2><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/09/Screen-Shot-2022-09-02-at-9.41.35-AM.png" class="kg-image" alt="Log analytics using ClickHouse" title="Elasticsearch vs ClickHouse"></figure><p>Logs are supposed to be immutable by nature and ClickHouse works best with immutable data. We were able to migrate one of the critical and significant log-producing applications from Elasticsearch to a much smaller ClickHouse cluster.</p><p>CPU and memory consumption on the inserter side were reduced by eight times. Each Elasticsearch document which used 600 bytes, came down to 60 bytes per row in ClickHouse. This storage gain allowed us to store 100% of the events in a newer setup. On the query side, the 99th percentile of the query latency also improved drastically.</p><p>Elasticsearch is great for full-text search and ClickHouse is great for analytics.</p>]]></content:encoded></item><item><title><![CDATA[Cloudflare's abuse policies & approach]]></title><description><![CDATA[Cloudflare launched nearly twelve years ago. Over that time, our set of services has become much more complicated. With that complexity we have developed policies around how we handle abuse of different features Cloudflare provides]]></description><link>https://blog.cloudflare.com/cloudflares-abuse-policies-and-approach/</link><guid isPermaLink="false">630e9a205f22d3000b330259</guid><category><![CDATA[Abuse]]></category><category><![CDATA[Freedom of Speech]]></category><category><![CDATA[Legal]]></category><dc:creator><![CDATA[Matthew Prince]]></dc:creator><pubDate>Wed, 31 Aug 2022 13:00:00 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/The-Cloudflare-Blog.png" medium="image"/><content:encoded><![CDATA[<!--kg-card-begin: markdown--><img src="http://blog.cloudflare.com/content/images/2022/08/The-Cloudflare-Blog.png" alt="Cloudflare's abuse policies & approach"><p><em><small>This post is also available in <a href="http://blog.cloudflare.com/zh-cn/cloudflares-abuse-policies-and-approach-zh-cn/">简体中文</a>, <a href="http://blog.cloudflare.com/ja-jp/cloudflares-abuse-policies-and-approach-ja-jp/">日本語</a>, <a href="http://blog.cloudflare.com/fr-fr/cloudflares-abuse-policies-and-approach-fr-fr/">Français</a>, <a href="http://blog.cloudflare.com/de-de/cloudflares-abuse-policies-and-approach-de-de/">Deutsch</a>, <a href="http://blog.cloudflare.com/es-es/cloudflares-abuse-policies-and-approach-es-es/">Español</a>.</small></em></p>
<!--kg-card-end: markdown--><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/The-Cloudflare-Blog-1.png" class="kg-image" alt="Cloudflare's abuse policies & approach"></figure><p>Cloudflare launched nearly twelve years ago. We’ve grown to operate a network that spans more than 275 cities in over 100 countries. We have millions of customers: from small businesses and individual developers to approximately 30 percent of the Fortune 500. Today, more than 20 percent of the web relies directly on Cloudflare’s services.</p><p>Over the time since we launched, our set of services has become much more complicated. With that complexity we have developed policies around how we handle abuse of different Cloudflare features. Just as a broad platform like Google has different abuse policies for search, Gmail, YouTube, and Blogger, Cloudflare has <a href="http://blog.cloudflare.com/out-of-the-clouds-and-into-the-weeds-cloudflares-approach-to-abuse-in-new-products/">developed different abuse policies</a> as we have introduced new products.</p><p>We published our updated approach to abuse last year at:</p><p><a href="https://www.cloudflare.com/trust-hub/abuse-approach/">https://www.cloudflare.com/trust-hub/abuse-approach/</a></p><p>However, as questions have arisen, we thought it made sense to describe those policies in more detail here.  </p><p>The policies we built reflect ideas and recommendations from human rights experts, activists, academics, and regulators. Our guiding principles require abuse policies to be specific to the service being used. This is to ensure that any actions we take both reflect the ability to address the harm and minimize unintended consequences. We believe that someone with an abuse complaint must have access to an abuse process to reach those who can most effectively and narrowly address their complaint — anonymously if necessary. And, critically, we strive always to be transparent about both our policies and the actions we take.</p><h3 id="cloudflare-s-products">Cloudflare's products</h3><p>Cloudflare provides a broad range of products that fall generally into three buckets: hosting products (e.g., Cloudflare Pages, Cloudflare Stream, Workers KV, Custom Error Pages), security services (e.g., DDoS Mitigation, Web Application Firewall, Cloudflare Access, Rate Limiting), and core Internet technology services (e.g., Authoritative DNS, Recursive DNS/1.1.1.1, WARP). For a complete list of our products and how they map to these categories, you can see our <a href="https://www.cloudflare.com/trust-hub/abuse-approach/">Abuse Hub</a>.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/pasted-image-0--2--1.png" class="kg-image" alt="Cloudflare's abuse policies & approach"></figure><p>As described below, our policies take a different approach on a product-by-product basis in each of these categories.</p><h3 id="hosting-products">Hosting products</h3><p>Hosting products are those products where Cloudflare is the ultimate host of the content. This is different from products where we are merely providing security or temporary caching services and the content is hosted elsewhere. Although many people confuse our security products with hosting services, we have distinctly different policies for each. Because the vast majority of Cloudflare customers do not yet use our hosting products, abuse complaints and actions involving these products are currently relatively rare. </p><p>Our decision to disable access to content in hosting products fundamentally results in that content being taken offline, at least until it is republished elsewhere. Hosting products are subject to our <a href="https://www.cloudflare.com/trust-hub/abuse-approach/">Acceptable Hosting Policy</a>. Under that policy, for these products, we may remove or disable access to content that we believe:</p><ul><li>Contains, displays, distributes, or encourages the creation of child sexual abuse material, or otherwise exploits or promotes the exploitation of minors.</li><li>Infringes on intellectual property rights.</li><li>Has been determined by appropriate legal process to be defamatory or libelous.</li><li>Engages in the unlawful distribution of controlled substances.</li><li>Facilitates human trafficking or prostitution in violation of the law.</li><li>Contains, installs, or disseminates any active malware, or uses our platform for exploit delivery (such as part of a command and control system).</li><li>Is otherwise illegal, harmful, or violates the rights of others, including content that discloses sensitive personal information, incites or exploits violence against people or animals, or seeks to defraud the public.</li></ul><p>We maintain discretion in how our Acceptable Hosting Policy is enforced, and generally seek to apply content restrictions as narrowly as possible. For instance, if a shopping cart platform with millions of customers uses Cloudflare Workers KV and one of their customers violates our Acceptable Hosting Policy, we will not automatically terminate the use of Cloudflare Workers KV for the entire platform.</p><p>Our guiding principle is that organizations closest to content are best at determining when the content is abusive. It also recognizes that overbroad takedowns can have significant unintended impact on access to content online.</p><h3 id="security-services">Security services</h3><p>The overwhelming majority of Cloudflare's millions of customers use only our security services. Cloudflare made a decision early in our history that we wanted to make security tools as widely available as possible. This meant that we provided many tools for free, or at minimal cost, to best limit the impact and effectiveness of a wide range of cyberattacks. Most of our customers pay us nothing.</p><p>Giving everyone the ability to sign up for our services online also reflects our view that cyberattacks not only should not be used for silencing vulnerable groups, but are not the appropriate mechanism for addressing problematic content online. We believe cyberattacks, in any form, should be relegated to the dustbin of history.</p><p>The decision to provide security tools so widely has meant that we've had to think carefully about when, or if, we ever terminate access to those services. We recognized that we needed to think through what the effect of a termination would be, and whether there was any way to set standards that could be applied in a fair, transparent and non-discriminatory way, consistent with human rights principles.</p><p>This is true not just for the content where a complaint may be filed  but also for the precedent the takedown sets. Our conclusion — informed by all of the many conversations we have had and the thoughtful discussion in the broader community — is that voluntarily terminating access to services that protect against cyberattack is not the correct approach.</p><h3 id="avoiding-an-abuse-of-power">Avoiding an abuse of power</h3><p>Some argue that we should terminate these services to content we find reprehensible so that others can launch attacks to knock it offline. That is the equivalent argument in the physical world that the fire department shouldn't respond to fires in the homes of people who do not possess sufficient moral character. Both in the physical world and online, that is a dangerous precedent, and one that is over the long term most likely to disproportionately harm vulnerable and marginalized communities.</p><p>Today, more than 20 percent of the web uses Cloudflare's security services. When considering our policies we need to be mindful of the impact we have and precedent we set for the Internet as a whole. Terminating security services for content that our team personally feels is disgusting and immoral would be the popular choice. But, in the long term, such choices make it more difficult to protect content that supports oppressed and marginalized voices against attacks.</p><h3 id="refining-our-policy-based-on-what-we-ve-learned">Refining our policy based on what we’ve learned</h3><p>This isn't hypothetical. Thousands of times per day we receive calls that we terminate security services based on content that someone reports as offensive. Most of these don’t make news. Most of the time these decisions don’t conflict with our moral views. Yet two times in the past we decided to terminate content from our security services because we found it reprehensible. In 2017, we terminated the neo-Nazi troll site <a href="http://blog.cloudflare.com/why-we-terminated-daily-stormer/">The Daily Stormer</a>. And in 2019, we terminated the conspiracy theory forum <a href="http://blog.cloudflare.com/terminating-service-for-8chan/">8chan</a>.</p><p>In a deeply troubling response, after both terminations we saw a dramatic increase in authoritarian regimes attempting to have us terminate security services for human rights organizations — often citing the language from our own justification back to us.</p><p>Since those decisions, we have had significant discussions with policy makers worldwide. From those discussions we concluded that the power to terminate security services for the sites was not a power Cloudflare should hold. Not because the content of those sites wasn't abhorrent — it was — but because security services most closely resemble Internet utilities.</p><p>Just as the telephone company doesn't terminate your line if you say awful, racist, bigoted things, we have concluded in consultation with politicians, policy makers, and experts that turning off security services because we think what you publish is despicable is the wrong policy. To be clear, just because we did it in a limited set of cases before doesn’t mean we were right when we did. Or that we will ever do it again.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/pasted-image-0--4--3.png" class="kg-image" alt="Cloudflare's abuse policies & approach"></figure><p>But that doesn’t mean that Cloudflare can’t play an important role in protecting those targeted by others on the Internet. We have long supported human rights groups, journalists, and other uniquely vulnerable entities online through <a href="https://www.cloudflare.com/galileo/">Project Galileo</a>. Project Galileo offers free cybersecurity services to nonprofits and advocacy groups that help strengthen our communities.</p><p>Through the <a href="https://www.cloudflare.com/athenian/">Athenian Project</a>, we also play a role in protecting election systems throughout the United States and abroad. Elections are one of the areas where the systems that administer them need to be fundamentally trustworthy and neutral. Making choices on what content is deserving or not of security services, especially in any way that could in any way be interpreted as political, would undermine our ability to provide trustworthy protection of election infrastructure.</p><h3 id="regulatory-realities">Regulatory realities</h3><p>Our policies also respond to regulatory realities. Internet content regulation laws passed over the last five years around the world have largely drawn a line between services that host content and those that provide security and conduit services. Even when these regulations impose obligations on platforms or hosts to moderate content, they exempt security and conduit services from playing the role of moderator without legal process. This is sensible regulation borne of a thorough regulatory process.</p><p>Our policies follow this well-considered regulatory guidance. We prevent security services from being used by sanctioned organizations and individuals. We also terminate security services for content which is illegal in the United States — where Cloudflare is headquartered. This includes Child Sexual Abuse Material (CSAM) as well as content subject to Fight Online Sex Trafficking Act (FOSTA). But, otherwise, we believe that cyberattacks are something that everyone should be free of. Even if we fundamentally disagree with the content.</p><p>In respect of the rule of law and due process, we follow legal process controlling security services. We will restrict content in geographies where we have received legal orders to do so. For instance, if a court in a country prohibits access to certain content, then, following that court's order, we generally will restrict access to that content in that country. That, in many cases, will limit the ability for the content to be accessed in the country. However, we recognize that just because content is illegal in one jurisdiction does not make it illegal in another, so we narrowly tailor these restrictions to align with the jurisdiction of the court or legal authority.</p><p>While we follow legal process, we also believe that transparency is critically important. To that end, wherever these content restrictions are imposed, we attempt to link to the particular legal order that required the content be restricted. This transparency is necessary for people to participate in the legal and legislative process. We find it deeply troubling when ISPs comply with court orders by invisibly blackholing content — not giving those who try to access it any idea of what legal regime prohibits it. Speech can be curtailed by law, but proper application of the Rule of Law requires whoever curtails it to be transparent about why they have.</p><h3 id="core-internet-technology-services">Core Internet technology services</h3><p>While we will generally follow legal orders to restrict security and conduit services, we have a higher bar for core Internet technology services like Authoritative DNS, Recursive DNS/1.1.1.1, and WARP. The challenge with these services is that restrictions on them are global in nature. You cannot easily restrict them just in one jurisdiction so the most restrictive law ends up applying globally.</p><p>We have generally challenged or appealed legal orders that attempt to restrict access to these core Internet technology services, even when a ruling only applies to our free customers. In doing so, we attempt to suggest to regulators or courts more tailored ways to restrict the content they may be concerned about.</p><p>Unfortunately, these cases are becoming more common where largely copyright holders are attempting to get a ruling in one jurisdiction and have it apply worldwide to terminate core Internet technology services and effectively wipe content offline. Again, we believe this is a dangerous precedent to set, placing the control of what content is allowed online in the hands of whatever jurisdiction is willing to be the most restrictive.</p><p>So far, we’ve largely been successful in making arguments that this is not the right way to regulate the Internet and getting these cases overturned. Holding this line we believe is fundamental for the healthy operation of the global Internet. But each showing of discretion across our security or core Internet technology services weakens our argument in these important cases.</p><h3 id="paying-versus-free">Paying versus free</h3><p>Cloudflare provides both free and paid services across all the categories above. Again, the majority of our customers use our free services and pay us nothing.</p><p>Although most of the concerns we see in our abuse process relate to our free customers, we do not have different moderation policies based on whether a customer is free versus paid. We do, however, believe that in cases where our values are diametrically opposed to a paying customer that we should take further steps to not only not profit from the customer, but to use any proceeds to further our companies’ values and oppose theirs. </p><p>For instance, when a site that opposed LGBTQ+ rights signed up for a paid version of DDoS mitigation service we worked with our Proudflare employee resource group to identify an organization that supported LGBTQ+ rights and donate 100 percent of the fees for our services to them. We don't and won't talk about these efforts publicly because we don't do them for marketing purposes; we do them because they are aligned with what we believe is morally correct.</p><h3 id="rule-of-law">Rule of Law</h3><p>While we believe we have an obligation to restrict the content that we host ourselves, we do not believe we have the political legitimacy to determine generally what is and is not online by restricting security or core Internet services. If that content is harmful, the right place to restrict it is legislatively.</p><p>We also believe that an Internet where cyberattacks are used to silence what's online is a broken Internet, no matter how much we may have empathy for the ends. As such, we will look to legal process, not popular opinion, to guide our decisions about when to terminate our security services or our core Internet technology services.</p><p>In spite what some may claim, we are not free speech absolutists. We do, however, believe in the Rule of Law. Different countries and jurisdictions around the world will determine what content is and is not allowed based on their own norms and laws. In assessing our obligations, we look to whether those laws are limited to the jurisdiction and consistent with our obligations to respect human rights under the <a href="https://www.ohchr.org/sites/default/files/documents/publications/guidingprinciplesbusinesshr_en.pdf">United Nations Guiding Principles on Business and Human Rights</a>.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/pasted-image-0--3--2.png" class="kg-image" alt="Cloudflare's abuse policies & approach"></figure><p>There remain many injustices in the world, and unfortunately much content online that we find reprehensible. We can solve some of these injustices, but we cannot solve them all. But, in the process of working to improve the security and functioning of the Internet, we need to make sure we don’t cause it long-term harm.</p><p>We will continue to have conversations about these challenges, and how best to approach securing the global Internet from cyberattack. We will also continue to cooperate with legitimate law enforcement to help investigate crimes, to <a href="https://www.cloudflare.com/galileo/">donate funds and services</a> to support equality, human rights, and other causes we believe in, and to participate in policy making around the world to help preserve the free and open Internet.</p>]]></content:encoded></item><item><title><![CDATA[Introducing thresholds in Security Event Alerting: a z-score love story]]></title><description><![CDATA[Today we are excited to announce thresholds for our Security Event Alerts: a new and improved way of detecting anomalous spikes of security events on your Internet properties. By introducing a threshold, we are able to make alerts more accurate and only notify you when it truly matters]]></description><link>https://blog.cloudflare.com/introducing-thresholds-in-security-event-alerting-a-z-score-love-story/</link><guid isPermaLink="false">63083d115f22d3000b3300a0</guid><category><![CDATA[Notifications]]></category><category><![CDATA[Alerts]]></category><category><![CDATA[WAF]]></category><category><![CDATA[Product News]]></category><category><![CDATA[Security]]></category><dc:creator><![CDATA[Kristina Galicova]]></dc:creator><pubDate>Tue, 30 Aug 2022 14:00:00 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/pasted-image-0--1--2.png" medium="image"/><content:encoded><![CDATA[<figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/pasted-image-0--1--1.png" class="kg-image" alt="Introducing thresholds in Security Event Alerting: a z-score love story"></figure><img src="http://blog.cloudflare.com/content/images/2022/08/pasted-image-0--1--2.png" alt="Introducing thresholds in Security Event Alerting: a z-score love story"><p>Today we are excited to announce thresholds for our Security Event Alerts: a new and improved way of detecting anomalous spikes of security events on your Internet properties. Previously, our calculations were based on z-score methodology alone, which was able to determine most of the significant spikes. By introducing a threshold, we are able to make alerts more accurate and only notify you when it truly matters. One can think of it as a romance between the two strategies. This is the story of how they met.</p><p>Author’s note: as an intern at Cloudflare I got to work on this project from start to finish from investigation all the way to the final product.</p><h3 id="once-upon-a-time">Once upon a time</h3><p>In the beginning, there were Security Event Alerts. Security Event Alerts are notifications that are sent whenever we detect a threat to your Internet property. As the name suggests, they track the number of security events, which are requests to your application that match security rules. For example, you can configure a security rule that blocks access from certain countries. Every time a user from that country tries to access your Internet property, it will log as a security event. While a security event may be harmless and fired as a result of the natural flow of traffic, it is important to alert on instances when a rule is fired more times than usual. Anomalous spikes of too many security events in a short period of time can indicate an attack. To find these anomalies and distinguish between the natural number of security events and that which poses a threat, we need a good strategy.</p><h3 id="the-lonely-life-of-a-z-score">The lonely life of a z-score</h3><p>Before a threshold entered the picture, our strategy worked only<em> </em>on the basis of a <a href="https://en.wikipedia.org/wiki/Standard_score">z-score</a>. Z-score is a methodology that looks at the number of standard deviations a certain data point is from the mean. In our current configuration, if a spike crosses the z-score value of 3.5, we send you an alert. This value was decided on after careful analysis of our customers’ data, finding it the most effective in determining a legitimate alert. Any lower and notifications will get noisy for smaller spikes. Any higher and we may miss out on significant events. You can read more about our z-score methodology in this <a href="http://blog.cloudflare.com/get-notified-when-your-site-is-under-attack/">blog post</a>. </p><p>The following graphs are an example of how the z-score method works. The first graph shows the number of security events over time, with a recent spike.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image3-9.png" class="kg-image" alt="Introducing thresholds in Security Event Alerting: a z-score love story" title="Chart"></figure><p>To determine whether this spike is significant, we calculate the z-score and check if the value is above 3.5:</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image4-3.png" class="kg-image" alt="Introducing thresholds in Security Event Alerting: a z-score love story" title="Chart"></figure><p>As the graph shows, the deviation is above 3.5 and so an alert is triggered.</p><p>However, relying on z-score becomes tricky for domains that experience no security events for a long period of time. With many security events at zero, the mean and standard deviation depress to zero as well. When a non-zero value finally appears, it will always be infinite standard deviations away from the mean. As a result, it will always trigger an alert even on spikes that do not pose any threat to your domain, such as the below:<br></p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image8.png" class="kg-image" alt="Introducing thresholds in Security Event Alerting: a z-score love story" title="Chart"></figure><p>With five security events, you are likely going to ignore this spike, as it is too low to indicate a meaningful threat. However, the z-score in this instance will be infinite:</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image5-3.png" class="kg-image" alt="Introducing thresholds in Security Event Alerting: a z-score love story" title="Chart"></figure><p>Since a z-score of infinity is greater than 3.5, an alert will be triggered. This means that customers with few security events would often be overwhelmed by event alerts that are not worth worrying about.</p><h3 id="letting-go-of-zeros">Letting go of zeros</h3><p>To avoid the mean and standard deviation becoming zero and thus alerting on every non-zero spike, zero values can be ignored in the calculation. In other words, to calculate the mean and standard deviation, only data points that are higher than zero will be considered. </p><p>With those conditions, the same spike to five security events will now generate a different z-score:</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image7.png" class="kg-image" alt="Introducing thresholds in Security Event Alerting: a z-score love story" title="Chart"></figure><p>Great! With the z-score at zero, it will no longer trigger an alert on the harmless spike! </p><p>But what about spikes that could be harmful? When calculations ignore zeros, we need enough non-zero data points to accurately determine the mean and standard deviation. If only one non-zero value is present, that data point determines the mean and standard deviation. As such, the mean will always be equal to the spike, z-score will always be zero and an alert will never be triggered:</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image6-1.png" class="kg-image" alt="Introducing thresholds in Security Event Alerting: a z-score love story" title="Chart"></figure><p>For a spike of 1000 events, we can tell that there is something wrong and we should trigger an alert. However, because there is only one non-zero data point, the z-score will remain zero: <br></p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image7-1.png" class="kg-image" alt="Introducing thresholds in Security Event Alerting: a z-score love story" title="Chart"></figure><p>The z-score does not cross the value 3.5 and an alert will not be triggered. </p><p>So what’s better? Including zeros in our calculations can skew the results for domains with too many zero events and alert them every time a spike appears. Not including zeros is mathematically wrong and will never alert on these spikes.</p><h3 id="threshold-the-prince-charming">Threshold, the prince charming</h3><p>Clearly, a z-score is not enough on its own. </p><p>Instead, we paired up the z-score with a threshold. The threshold represents the raw number of security events an Internet property can have, below which an alert will not be sent. While z-score checks whether the spike is at least 3.5 standard deviations above the mean, the threshold makes sure it is above a certain static value. If both of these conditions are met, we will send you an alert: <br></p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image1-17.png" class="kg-image" alt="Introducing thresholds in Security Event Alerting: a z-score love story" title="Chart"></figure><p>The above spike crosses the threshold of 200 security events. We now have to check that the z-score is above 3.5:</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image9.png" class="kg-image" alt="Introducing thresholds in Security Event Alerting: a z-score love story" title="Chart"></figure><p>The z-score value crosses 3.5 and an alert will be sent.</p><p>A threshold for the number of security events comes as the perfect complement. By itself, the threshold cannot determine whether something is a spike, and would simply alert on any value crossing it. This <a href="http://blog.cloudflare.com/smarter-origin-service-level-monitoring/">blog post</a> describes in more detail why thresholds alone do not work. However, when paired with z-score, they are able to share their strengths and cover for each other's weaknesses. If the z-score falsely detects an insignificant spike, the threshold will stop the alert from triggering. Conversely, if a value does cross the security events threshold, the z-score ensures there is a reasonable variance from the data average before allowing an alert to be sent.</p><h3 id="the-invaluable-value">The invaluable value</h3><p>To foster a successful relationship between the z-score and security events threshold, we needed to determine the most effective threshold value. After careful analysis of our previous attacks on customers, we set the value to 200. This number is high enough to filter out the smaller, noisier spikes, but low enough to expose any threats.</p><h3 id="am-i-invited-to-the-wedding">Am I invited to the wedding?</h3><p>Yes, you are! The z-score and threshold relationship is already enabled for all WAF customers, so all you need to do is sit back and relax. For enterprise customers, the threshold will be applied to each type of alert enabled on your domain.</p><h3 id="happily-ever-after">Happily ever after</h3><p>The story certainly does not end here. We are constantly iterating on our alerts, so keep an eye out for future updates on the road to make our algorithms even more personalized for your Internet properties!</p>]]></content:encoded></item><item><title><![CDATA[Performance isolation in a multi-tenant database environment]]></title><description><![CDATA[Enforcing stricter performance isolation across neighboring tenants who rely on our storage infrastructure]]></description><link>https://blog.cloudflare.com/performance-isolation-in-a-multi-tenant-database-environment/</link><guid isPermaLink="false">63085d0b5f22d3000b33011f</guid><category><![CDATA[Edge Database]]></category><category><![CDATA[PgBouncer]]></category><category><![CDATA[PostgresSQL]]></category><category><![CDATA[Multi-Tenant]]></category><dc:creator><![CDATA[Justin Kwan]]></dc:creator><pubDate>Fri, 26 Aug 2022 15:08:06 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/1-1.png" medium="image"/><content:encoded><![CDATA[<figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/1.png" class="kg-image" alt="Performance isolation in a multi-tenant database environment"></figure><img src="http://blog.cloudflare.com/content/images/2022/08/1-1.png" alt="Performance isolation in a multi-tenant database environment"><p>Operating at Cloudflare scale means that across the technology stack we spend a great deal of time handling different load conditions. In this blog post we talk about how we solved performance difficulties with our Postgres clusters. These clusters support a large number of tenants and highly variable load conditions leading to the need to isolate activity to prevent tenants taking too much time from others. Welcome to real-world, large database cluster management!</p><p>As an intern at Cloudflare I got to work on improving how our database clusters behave under load and open source the resulting code.</p><p>Cloudflare operates production Postgres clusters across multiple regions in data centers. Some of our earliest service offerings, such as our DNS Resolver, Firewall, and DDoS Protection, depend on our Postgres clusters' high availability for OLTP workloads. The high availability cluster manager, <a href="https://github.com/sorintlab/stolon">Stolon</a>, is employed across all clusters to independently control and replicate data across Postgres instances and elect Postgres leaders and failover under high load scenarios.</p><p>PgBouncer and HAProxy act as the gateway layer in each cluster. Each tenant acquires client-side connections from PgBouncer instead of Postgres directly. PgBouncer holds a pool of maximum server-side connections to Postgres, allocating those across multiple tenants to prevent Postgres connection starvation. From here, PgBouncer forwards queries to HAProxy, which load balances across Postgres primary and read replicas.</p><h2 id="problem">Problem</h2><p>Our multi-tenant Postgres instances operate on bare metal servers in non-containerized environments. Each backend application service is considered a single tenant, where they may use one of multiple Postgres roles. Due to each cluster serving multiple tenants, all tenants share and contend for available system resources such as CPU time, memory, disk IO on each cluster machine, as well as finite database resources such as server-side Postgres connections and table locks. Each tenant has a unique workload that varies in system level resource consumption, making it impossible to enforce throttling using a global value.</p><p>This has become problematic in production affecting neighboring tenants:</p><ul><li><strong>Throughput</strong>. A tenant may issue a burst of transactions, starving shared resources from other tenants and degrading their performance.</li><li><strong>Latency</strong>: A single tenant may issue very long or expensive queries, often concurrently, such as large table scans for ETL extraction or queries with lengthy table locks. </li></ul><p>Both of these scenarios can result in degraded query execution for neighboring tenants. Their transactions may hang or take significantly longer to execute (higher latency) due to either reduced CPU share time, or slower disk IO operations due to many seeks from misbehaving tenant(s). Moreover, other tenants may be blocked from acquiring database connections from the database proxy level (PgBouncer) due to existing ones being held during long and expensive queries.</p><h2 id="previous-solution">Previous solution</h2><p>When database cluster load significantly increases, finding which tenants are responsible is the first challenge. Some techniques include searching through all tenants' previous queries under typical system load and determining whether any new expensive queries have been introduced under the Postgres' pg_stat_activity view.</p><h3 id="database-concurrency-throttling">Database concurrency throttling</h3><p>Once the misbehaving tenants are identified, Postgres server-side connection limits are manually enforced using the Postgres query.</p><pre><code>ALTER USER "some_bad-user" WITH CONNECTION LIMIT 123;</code></pre><p>This essentially restricts or “squeezes” the concurrent throughput for a single user, where each tenant will only be able to exhaust their share of connections. </p><p>Manual concurrency (connection) throttling has shown improvements in shedding load in Postgres during high production workloads:</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/2.png" class="kg-image" alt="Performance isolation in a multi-tenant database environment"></figure><p>While we have seen success with this approach, it is not perfect and is horribly manual. It also suffers from the following:</p><ul><li>Postgres does not immediately kill existing tenant connections when a new user limit is set; the user may continue to issue bursty or expensive queries.</li><li>Tenants may still issue very expensive, resource intensive queries (affecting neighboring tenants) even if their concurrency (connection pool size) is reduced.</li><li>Manually applying connection limits against a misbehaving tenant is toil; an SRE could be paged to physically apply the new limit at any time of the day.</li><li>Manually analyzing and detecting misbehaving tenants based on queries can be time-consuming and stressful especially during an incident, requiring production SQL analysis experience.</li><li>Additionally, applying new throttling limits per user/pool, such as the allocated connection count, can be arbitrary and experimental while requiring extensive understanding of tenant workloads.</li><li>Oftentimes, Postgres may be under so much load that it begins to hang (CPU starvation). SREs may be unable to manually throttle tenants through native interfaces once a high load situation occurs.</li></ul><h2 id="new-solution">New solution</h2><h3 id="gateway-concurrency-throttling">Gateway concurrency throttling</h3><p>Typically, the system level resource consumption of a query is difficult to control and isolate once submitted to the server or database system for execution. However, a common approach is to intercept and throttle connections or queries at the gateway layer, controlling per user/pool traffic characteristics based on system resource consumption.</p><p>We have implemented connection throttling at our database proxy server/connection pooler, PgBouncer. Previously, PgBouncer’s user level connection limits would not kill existing connections, but only prevent exceeding it. We now support the ability to throttle and kill existing connections owned by each user or each user’s connection pool statically via configuration or at runtime via new administrative commands.</p><h5 id="pgbouncer-configuration">PgBouncer Configuration</h5><pre><code>[users]
dns_service_user = max_user_connections=60
firewall_service_user = max_user_connections=80
[pools]
user1.database1 = pool_size=90</code></pre><h5 id="pgbouncer-runtime-commands">PgBouncer Runtime Commands</h5><pre><code>SET USER dns_service_user = ‘max_user_connections=40’;
SET POOL dns_service_user.dns_db = ‘pool_size=30’;</code></pre><p>This required major bug fixes, refactoring and implementation work in our fork of PgBouncer. We’ve also raised multiple pull requests to contribute all of our features to PgBouncer open source. To read about all of our work in PgBouncer, read <a href="http://blog.cloudflare.com/open-sourcing-our-fork-of-pgbouncer/">this blog</a>.</p><p>These new features now allow for faster and more granular "load shedding" against a misbehaving tenant’s concurrency (connection pool, user and database pair), while enabling stricter performance isolation.</p><h2 id="future-solutions">Future solutions</h2><p>We are continuing to build infrastructure components that monitor per-tenant resource consumption and detect which tenants are misbehaving based on system resource indicators against historical baselines. We aim to automate connection and query throttling against tenants using these new administrative commands.</p><p>We are also experimenting with various automated approaches to enforce strict tenant performance isolation.</p><h3 id="congestion-avoidance">Congestion avoidance</h3><p>An adaptation of the TCP Vegas congestion avoidance algorithm can be employed to adaptively estimate and enforce each tenant’s optimal concurrency while still maintaining low latency and high throughput for neighboring tenants. This approach does not require resource consumption profiling, manual threshold tuning, knowledge of underlying system hardware, or expensive computation.</p><p>Traditionally, TCP Vegas converges to the initially unknown and optimal congestion window (max packets that can be sent concurrently). In the same spirit, we can treat the unknown congestion window as the optimal concurrency or connection pool size for database queries. At the gateway layer, PgBouncer, each tenant will begin with a small connection pool size, while we dynamically sample each tenant’s transaction’s round trip time (RTT) against Postgres. We gradually increase the connection pool size (congestion window) of a tenant so long as their transaction RTTs do not deteriorate.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/3.png" class="kg-image" alt="Performance isolation in a multi-tenant database environment"></figure><p>When a tenant's sampled transaction latency increases, the formula's minimum by sampled request latency ratio will decrease, naturally reducing the tenant's available concurrency which reduces database load.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/4.png" class="kg-image" alt="Performance isolation in a multi-tenant database environment"></figure><p>Essentially, this algorithm will "back off" when observing high query latencies as the indicator of high database load, regardless of whether the latency is due to CPU time or disk/network IO blocking, etc. This formula will converge to find the optimal concurrency limit (connection pool size) since the latency ratio always converges to 0 with sufficiently large sample request latencies. The square root of the current tenant pool size is chosen as a constant request "burst" headroom because of its fast growth and being relatively large for small pool sizes (when latencies are low) but converges when the pool size is reduced (when latencies are high).</p><p>Rather than reactively shedding load, congestion avoidance preventatively or “smoothly” <strong>throttles traffic before load induced performance degradation becomes an issue</strong>. This algorithm aims to prevent database server resource starvation which causes other queries to hang.</p><p>Theoretically, if one tenant misbehaves and causes load induced latency for others, this TCP congestion algorithm may incorrectly blindly throttle all tenants. Hence why it may be necessary to apply this adaptive throttling only against tenants with high CPU to latency correlation when the system performance is degrading.</p><h3 id="tenant-resource-quotas">Tenant resource quotas</h3><p>Configurable resource quotas can be introduced per each tenant. Upstream application service tenants are restricted to their allocated share of resources expressed as CPU % utilized per second and max memory. If a tenant overuses their share, the database gateway (PgBouncer) should throttle their concurrency, queries per second and ingress bytes to force consumption within their allocated slice.</p><p>Resource throttling a tenant must not "spillover" or affect other tenants accessing the same cluster. This could otherwise reduce the availability of other customer-facing applications and violate SLO (service-level objectives). Resource restriction must be isolated to each tenant.</p><p>If traffic is low against Postgres instances, tenants should be permitted to exceed their allocation limit. However, when load against the cluster degrades the entire performance of the system (latency), the tenant's limit must be re-enforced at the gateway layer, PgBouncer. We can make deductions around the health of the entire database server based on indicators such as average query latency’s rate of change against a predefined threshold. All tenants should agree that a surplus in resource consumption may result in query throttling of any pattern.</p><p>Each tenant has a <strong>unique and variable workload</strong>, which may degrade multi tenant performance at any time. Quick detection requires profiling the baseline resource consumption of each tenant’s (or tenant’s connection pooled) workload against each local Postgres server (backend pids) in near real-time. From here, we can correlate the “baseline” traffic characteristics with system level resource consumption per database instance.</p><p>Taking an average or generalizing statistical measures across distributed nodes (each tenant's resource consumption on Postgres instances in this case) can be inaccurate due to high variance in traffic against leader vs replica instances. This would lead to faulty throttling decisions applied against users. For instance, we should not throttle a user’s concurrency on an idle read replica even if the user consumes excessive resources on the primary database instance. It is preferable to capture tenant consumption on a per Postgres instance level, and enforce throttling per instance rather than across the entire cluster.</p><p>Multivariable regression can be employed to model the relationship between independent variables (concurrency, queries per second, ingested bytes) against the dependent variables (system level resource consumption). We can calculate and enforce the optimal independent variables per tenant under high load scenarios. To account for workload changes, regression<strong> adaptability vs accuracy</strong> will need to be tuned by adjusting the sliding window size (amount of time to retain profiled data points) when capturing workload consumption.</p><h3 id="gateway-query-queuing">Gateway query queuing</h3><p>User queries can be prioritized for submission to Postgres at the gateway layer (PgBouncer). Within a one or multiple global priority queues, query submissions by all tenants are ordered based on the current resource consumption of the tenant’s connection pool or the tenant itself. Alternatively, ordering can be based on each query’s historical resource consumption, where each query is independently profiled. Based on changes in tenant resource consumption captured from each Postgres instance’s server, all queued queries can be reordered every time the scheduler forwards a query to be submitted.</p><p>To prevent priority queue starvation (one tenant’s query is at the end of the queue and is never executed), the gateway level query queuing can be configured to only enable when there is peak load/traffic against the Postgres instance. Or, the time of enqueueing a query can be factored into the priority ordering.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/6.png" class="kg-image" alt="Performance isolation in a multi-tenant database environment"></figure><p>This approach would isolate tenant performance by allowing non-offending tenants to continue reserving connections and executing queries (such as critical health monitoring queries). <strong>Higher latency would only be observed from the tenants that are utilizing more resources</strong> (from many/expensive transactions). This approach is straightforward to understand, generic in application (can queue transactions based on other input metrics), and <strong>non-destructive</strong> as it does not kill client/server connections, and should only drop queries when the in-memory priority queue reaches capacity.</p><h3 id="conclusion">Conclusion</h3><p>Performance isolation in our multi-tenant storage environment continues to be a very interesting challenge that touches areas including OS resource management, database internals, queueing theory, congestion algorithms and even statistics. We’d love to hear how the community has tackled the “noisy neighbor” problem by isolating tenant performance at scale!</p>]]></content:encoded></item><item><title><![CDATA[Open sourcing our fork of PgBouncer]]></title><description><![CDATA[We are releasing our internal fork of PgBouncer, filled with authentication bug fixes and new features around per user and connection pool isolation]]></description><link>https://blog.cloudflare.com/open-sourcing-our-fork-of-pgbouncer/</link><guid isPermaLink="false">63083f8c5f22d3000b3300f2</guid><category><![CDATA[PgBouncer]]></category><category><![CDATA[PostgresSQL]]></category><category><![CDATA[Edge Database]]></category><category><![CDATA[Open Source]]></category><dc:creator><![CDATA[Justin Kwan]]></dc:creator><pubDate>Fri, 26 Aug 2022 14:30:37 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/Magic-Nat-1.png" medium="image"/><content:encoded><![CDATA[<figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/Magic-Nat.png" class="kg-image" alt="Open sourcing our fork of PgBouncer"></figure><img src="http://blog.cloudflare.com/content/images/2022/08/Magic-Nat-1.png" alt="Open sourcing our fork of PgBouncer"><p>Cloudflare operates highly available Postgres production clusters across multiple data centers, supporting the transactional workloads of our core service offerings such as our DNS Resolver, Firewall, and DDoS Protection.</p><p>Multiple PgBouncer instances sit at the front of the gateway layer per each cluster, acting as a TCP proxy that provides Postgres connection pooling. PgBouncer’s pooling enables upstream applications to connect to Postgres, without having to constantly open and close connections (expensive) at the database level, while also reducing the number of Postgres connections used. Each tenant acquires client-side connections from PgBouncer instead of Postgres directly.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/Frame-673.png" class="kg-image" alt="Open sourcing our fork of PgBouncer"></figure><p>PgBouncer will hold a pool of maximum server-side connections to Postgres, allocating those across multiple tenants to prevent Postgres connection starvation. From here, PgBouncer will forward backend queries to HAProxy, which load balances across Postgres primary and read replicas.</p><p>As an intern at Cloudflare I got to work on improving how our database clusters behave under load and open source the resulting code.</p><p>We run our Postgres infrastructure in non-containerized, bare metal environments which consequently leads to multitenant resource contention between Postgres users. To enforce stricter tenant performance isolation at the database level (CPU time utilized, memory consumption, disk IO operations), we’d like to configure and enforce connection limits per user and connection pool at PgBouncer.</p><p>To do that we had to add features and fix bugs in PgBouncer. Rather than continue to maintain a private fork we are open sourcing our code for others to use.</p><h3 id="authentication-rejection">Authentication Rejection</h3><p>The PgBouncer connection pooler offers options to enforce server connection pool size limits (effective concurrency) per user via static configuration. However, an authentication bug upstream prevented these features from correctly working when Postgres was set to use HBA authentication. Administrators who sensibly use server-side authentication could not take advantage of these user-level features.</p><p>This ongoing issue has also been experienced by others in the open-source community:</p><p><a href="https://github.com/pgbouncer/pgbouncer/issues/484">https://github.com/pgbouncer/pgbouncer/issues/484</a><br><a href="https://github.com/pgbouncer/pgbouncer/issues/596">https://github.com/pgbouncer/pgbouncer/issues/596</a></p><h3 id="root-cause">Root Cause</h3><p>PgBouncer needs a Postgres user’s password when proxying submitted queries from client connection to a Postgres server connection. PgBouncer will fetch a user’s Postgres password defined in userlist.txt (auth_file) when a user first logs in to compare against the provided password. However, if the user is not defined in userlist.txt, Pgbouncer will fetch their password from the Postgres <a href="https://www.postgresql.org/docs/current/view-pg-shadow.html">pg_shadow</a> system view for comparison. This password will be used when PgBouncer subsequently forwards queries from this user to Postgres. The same applies when Postgres is configured to use HBA authentication.</p><p>Following serious debugging efforts and time spent in GDB, we found that multiple user objects are typically created for a single real user: via configuration loading from the [users] section and upon the user’s first login. In PgBouncer, any users requiring a shadow auth query would be stored under their respective database struct instance, whereas any user with a password defined in userlist.txt would be stored globally. Because the non-authenticated user already existed in memory after being parsed from the [users] section, PgBouncer assumed that the user was defined in userlist.txt, where the shadow authentication query could be skipped. It would not bother to fetch and set the user’s password upon first login, resulting in an empty user password. This is why subsequent queries submitted by the user would be rejected with authentication failure at Postgres.</p><p>To solve this, we simplified the code to globally store all users in one place rather than store different types of users (requiring different methods of authentication) in a disaggregated fashion per database or globally. Also, rather than assuming a user is authenticated if they merely exist, we keep track of whether the user requires authentication via auth query or from fetching their password from userlist.txt. This depends on how they were created.</p><p>We saw the value in troubleshooting and fixing these issues; it would unlock an entire class of features in PgBouncer for our use cases, while benefiting many in the open-source community.</p><h3 id="new-features">New Features</h3><p>We’ve also done work to implement and support additional features in PgBouncer to enforce stricter tenant performance isolation.</p><p>Previously, PgBouncer would only prevent tenants from exceeding preconfigured limits, not particularly helpful when it’s too late and a user is misbehaving or already has too many connections. PgBouncer now supports enforcing or shrinking per user connection pool limits at runtime, <strong>where it is most critically needed</strong> to throttle tenants who are issuing a burst of expensive queries, or are hogging connections from other tenants. We’ve also implemented new administrative commands to throttle the maximum connections per user or per pool at runtime.</p><p>PgBouncer also now supports statically configuring and dynamically enforcing connection limits per connection pool. This feature is extremely important in order to granularly throttle a tenant’s misbehaving connection pool without throttling and reducing availability on its other non-misbehaving pools.</p><h5 id="pgbouncer-configuration">PgBouncer Configuration<br></h5><pre><code>[users]
dns_service_user = max_user_connections=60
firewall_service_user = max_user_connections=80
[pools]
user1.database1 = pool_size=90</code></pre><h5 id="pgbouncer-runtime-commands">PgBouncer Runtime Commands</h5><pre><code>SET USER dns_service_user = ‘max_user_connections=40’;
SET POOL dns_service_user.dns_db = ‘pool_size=30’;</code></pre><p>These new features required major refactoring around how PgBouncer stores users, databases weakly referenced and stored passwords of different users, and how we enforce killing server side connections while still in use.</p><h3 id="conclusion">Conclusion</h3><p>We are committed to improving PgBouncer in open source and contributing all of our features to benefit the wider community. If you are interested, please consider contributing to our <a href="https://github.com/cloudflare/cf-pgbouncer">open source PgBouncer fork</a>. After all, it is the community that makes PgBouncer possible!</p>]]></content:encoded></item><item><title><![CDATA[Deep dives & how the Internet works]]></title><description><![CDATA[We have amazing deep dives in our blog, but also research and how the Internet works kind of stories. Here are some highlights from 2022, and before (with glimpses of our history).]]></description><link>https://blog.cloudflare.com/deep-dives-how-the-internet-works/</link><guid isPermaLink="false">630794175f22d3000b32ffcd</guid><category><![CDATA[Deep Dive]]></category><category><![CDATA[Trends]]></category><category><![CDATA[Internet Traffic]]></category><category><![CDATA[Research]]></category><category><![CDATA[Reading List]]></category><dc:creator><![CDATA[João Tomé]]></dc:creator><pubDate>Thu, 25 Aug 2022 18:08:00 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/pasted-image-0-2.png" medium="image"/><content:encoded><![CDATA[<figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/pasted-image-0-1.png" class="kg-image" alt="Deep dives & how the Internet works"></figure><img src="http://blog.cloudflare.com/content/images/2022/08/pasted-image-0-2.png" alt="Deep dives & how the Internet works"><p>When August comes, for many, at least in the Northern Hemisphere, it’s time to enjoy summer and/or vacations. Here are some deep dive reading suggestions from our Cloudflare Blog for any time, weather or time of the year. There’s also some reading material on how the Internet works, and a glimpse into our history. </p><p>To create the list (that goes beyond 2022), initially we asked inside the company for favorite blog posts. Many explained how a particular blog post made them want to work at Cloudflare (including some of those who have been at the company for many years). And then, we also heard from readers by asking the question on <a href="https://twitter.com/Cloudflare/status/1555557749307719680">our Twitter</a> account: “What’s your favorite blog post from the Cloudflare Blog and why?” </p><h3 id="2022-deep-dive-trends-odyssey">2022, deep dive &amp; trends odyssey </h3><p>In early July (thinking of the July 4 US holiday) we did a <a href="http://blog.cloudflare.com/july-4-2022-reading-list/">sum up</a> where some of the more recent blog posts were referenced. We’ve added a few to that list: </p><ul><li><strong>Eliminating CAPTCHAs on iPhones and Macs (</strong><a href="http://blog.cloudflare.com/eliminating-captchas-on-iphones-and-macs-using-new-standard/"><strong>✍️</strong></a><strong>)</strong> <br><a href="http://blog.cloudflare.com/eliminating-captchas-on-iphones-and-macs-using-new-standard/"><strong>How it works</strong></a> using open standards. On this topic, you can also read the detailed blog post from our research team, from <a href="http://blog.cloudflare.com/introducing-cryptographic-attestation-of-personhood">2021</a>: <em>Humanity wastes about 500 years per day on CAPTCHAs. It’s time to end this madness</em>. <br><br></li><li><strong>Optimizing TCP for high WAN throughput while preserving low latency</strong> <strong>(</strong><a href="http://blog.cloudflare.com/optimizing-tcp-for-high-throughput-and-low-latency/"><strong>✍️</strong></a><strong>)</strong> <strong> </strong> <br>If you like networks, <a href="http://blog.cloudflare.com/optimizing-tcp-for-high-throughput-and-low-latency/"><strong>this</strong></a> is an in depth look of how we tune TCP parameters for low latency and high throughput.<br><br></li><li><strong>Live-patching the Linux kernel (</strong><a href="http://blog.cloudflare.com/live-patch-security-vulnerabilities-with-ebpf-lsm/"><strong>✍️</strong></a><strong>)</strong> <strong> </strong><br>A detail focused <a href="http://blog.cloudflare.com/live-patch-security-vulnerabilities-with-ebpf-lsm/"><strong>blog</strong></a> focused on using eBPF. Code, Makefiles and more within.<br><br></li><li><strong>Early Hints in the real world (</strong><a href="http://blog.cloudflare.com/early-hints-performance/"><strong>✍️</strong></a><strong>)</strong>  <br><a href="http://blog.cloudflare.com/early-hints-performance/"><strong>In depth data</strong></a><strong> </strong>about it where we show how much faster the web is with it (in a Cloudflare, Google, and Shopify partnership).<br><br></li><li><strong>Internet Explorer, we hardly knew ye (</strong><a href="http://blog.cloudflare.com/internet-explorer-retired/"><strong>✍️</strong></a><strong>)</strong> <strong> </strong><br>A look at the demise of <a href="http://blog.cloudflare.com/internet-explorer-retired/"><strong>Internet Explorer</strong></a> and the rise of the Edge browser (after Microsoft announced the end-of-life for IE).<br><br></li><li><strong>When the window is not fully open, your TCP stack is doing more than you think (</strong><a href="http://blog.cloudflare.com/when-the-window-is-not-fully-open-your-tcp-stack-is-doing-more-than-you-think/"><strong>✍️</strong></a><strong>)</strong> <strong> </strong><br>A recent deep dive <a href="http://blog.cloudflare.com/when-the-window-is-not-fully-open-your-tcp-stack-is-doing-more-than-you-think/"><strong>shows</strong></a> how Linux manages TCP receive buffers and windows, and how to tune the TCP connection for the best speed. Similar blogs are: <a href="http://blog.cloudflare.com/how-to-stop-running-out-of-ephemeral-ports-and-start-to-love-long-lived-connections/"><em>How</em></a><em> to stop running out of ephemeral ports and start to love long-lived connections</em>; <a href="http://blog.cloudflare.com/everything-you-ever-wanted-to-know-about-udp-sockets-but-were-afraid-to-ask-part-1/"><em>Everything</em></a><em> you ever wanted to know about UDP sockets but were afraid to ask</em>.<br><br></li><li><strong>How Ramadan shows up in Internet trends (</strong><a href="http://blog.cloudflare.com/how-ramadan-shows-up-in-internet-trends/"><strong>✍️</strong></a><strong>)</strong><br>What happens to the Internet traffic in countries where many observe <a href="http://blog.cloudflare.com/how-ramadan-shows-up-in-internet-trends/"><strong>Ramadan</strong></a>? Depending on the country, there are clear shifts and changing patterns in Internet use, particularly before dawn and after sunset. This is all coming from our <a href="http://blog.cloudflare.com/tag/cloudflare-radar/">Radar</a> platform. We can see many <a href="http://blog.cloudflare.com/tag/cloudflare-radar/">human trends</a>, from a relevant <a href="http://blog.cloudflare.com/cloudflares-view-of-the-rogers-communications-outage-in-canada/">outage</a> in a country (here’s the list of <a href="http://blog.cloudflare.com/q2-2022-internet-disruption-summary/">Q2 2022 disruptions</a>), to events like <a href="http://blog.cloudflare.com/french-elections-2022-runoff/">elections</a>, the <a href="http://blog.cloudflare.com/eurovision-2022-internet-trends/">Eurovision</a>, the ‘<a href="http://blog.cloudflare.com/queens-platinum-jubilee/">Jubilee</a>’ celebration or the <a href="http://blog.cloudflare.com/how-the-james-webb-telescopes-cosmic-pictures-impacted-the-internet/">James Webb Telescope</a> pictures revelation. </li></ul><p></p><h3 id="2022-research-focused">2022, research focused</h3><ul><li><strong>Hertzbleed attack (</strong><a href="http://blog.cloudflare.com/hertzbleed-explained/"><strong>✍️</strong></a><strong>)</strong> <strong>  </strong><br>A <a href="http://blog.cloudflare.com/hertzbleed-explained/"><strong>deep explainer</strong></a><strong> </strong>where we compare a runner in a long distance race with how CPU frequency scaling leads to a nasty side channel affecting cryptographic algorithms. Don’t be confused with the older and impactful <a href="http://blog.cloudflare.com/heartbleed-revisited/">Heartbleed</a>.<br><br></li><li><strong><strong><strong>Future-proofing SaltStack (</strong><a href="http://blog.cloudflare.com/future-proofing-saltstack/"><strong>✍️</strong></a><strong>)</strong> <strong>  </strong></strong></strong><br>A <a href="http://blog.cloudflare.com/future-proofing-saltstack/"><strong>chronicle</strong></a> of our path of making the <a href="https://saltproject.io/">SaltStack</a> system quantum-secure. In an extra post-quantum blog post, we highlight how we are <a href="http://blog.cloudflare.com/post-quantumify-cloudflare/">preparing the Internet and our infrastructure</a> for the arrival of quantum computers.<br><br></li><li><strong>Unlocking QUIC’s proxying potential with MASQUE (</strong><a href="http://blog.cloudflare.com/unlocking-quic-proxying-potential/"><strong>✍️</strong></a><strong>)</strong> <br>A <a href="http://blog.cloudflare.com/unlocking-quic-proxying-potential/"><strong>deep dive</strong></a> into <a href="http://blog.cloudflare.com/tag/quic/">QUIC</a> transport protocol and a good up to date way to know more about it (related: <a href="http://blog.cloudflare.com/cloudflare-view-http3-usage/">HTTP usage trends</a>).<br><br></li><li><strong>HPKE: Standardizing public-key encryption (finally!)</strong> <strong>(</strong><a href="http://blog.cloudflare.com/hybrid-public-key-encryption/"><strong>✍️</strong></a><strong>)</strong>  <br>Two research groups have <a href="http://blog.cloudflare.com/hybrid-public-key-encryption/"><strong>finally published</strong></a> the next reusable, and future-proof generation of (hybrid) public-key encryption (PKE) for Internet protocols and applications: Hybrid Public Key Encryption (HPKE).<br><br></li><li><strong>Sizing Up Post-Quantum Signatures (</strong><a href="http://blog.cloudflare.com/sizing-up-post-quantum-signatures/"><strong>✍️</strong></a><strong>)</strong>  <br>This <a href="http://blog.cloudflare.com/sizing-up-post-quantum-signatures/"><strong>blog</strong></a> (followed by this deep dive <a href="http://blog.cloudflare.com/post-quantum-signatures/">one</a> that includes quotes from Ancient Greece) was highlighted by a reader as “life changing”. It shows the peculiar relationship between PQC (post-quantum cryptography) signatures and <a href="https://www.cloudflare.com/learning/ssl/transport-layer-security-tls/">TLS</a> (Transport Layer Security) size and connection quality. It’s research about how quantum computers could unlock the next age of innovation, and will break the majority of the cryptography used to protect our web browsing (more on that below). But it is also about how to make a website really fast.</li></ul><p>If you like Twitter threads, <a href="https://twitter.com/grittygrease/status/1552391305405386752"><strong>here</strong></a> is a recent one from our Head of Cloudflare Research, <a href="https://research.cloudflare.com/people/nick-sullivan/">Nick Sullivan</a>, that explains in simple terms the way privacy on the Internet works and challenges in protecting it now and for the future. </p><p>This month we also did a <a href="http://blog.cloudflare.com/2022-attacks-an-august-reading-list-to-go-shields-up/">full reading list/guide</a> with our blog posts about all sorts of <strong>attacks</strong> (from DDoS to phishing, malware or ransomware) and how to stay protected in 2022.</p><h3 id="how-does-it-the-internet-work">How does it (the Internet) work</h3><ul><li><strong>Cloudflare’s view of the Rogers Communications outage in Canada (</strong><a href="http://blog.cloudflare.com/cloudflares-view-of-the-rogers-communications-outage-in-canada/"><strong>✍️</strong></a><strong> 2022)</strong> <br>One of the largest ISPs in Canada, Rogers Communications, had a huge outage on July 8, 2022, that lasted for more than 17 hours. From our view of the Internet, <a href="http://blog.cloudflare.com/cloudflares-view-of-the-rogers-communications-outage-in-canada/"><strong>we show</strong></a> why we concluded it seemed caused by an internal error and how the Internet, being a network of networks, all bound together by <a href="https://www.cloudflare.com/learning/security/glossary/what-is-bgp/">BGP</a>, was related to the disruption.<br><br></li><li><strong><strong><strong>Understanding how Facebook disappeared from the Internet (</strong><a href="http://blog.cloudflare.com/october-2021-facebook-outage/"><strong>✍️</strong></a><strong> 2021). </strong></strong></strong><br>“Facebook can't be down, can it?”, we thought, for a second, on October 4, 2021. It was, and we had a <a href="http://blog.cloudflare.com/october-2021-facebook-outage/"><strong>deep dive</strong></a> about it, where BGP was also ‘king’. </li></ul><p><em>Albert Einstein's </em><a href="https://en.wikipedia.org/wiki/Special_relativity"><em>special theory of relativity</em></a><em> famously dictates that no known object can travel faster than the speed of light in vacuum, which is 299,792 km/s. </em></p><ul><li><strong>Welcome to Speed Week and a Waitless Internet</strong> <strong>(</strong><a href="http://blog.cloudflare.com/fastest-internet/"><strong>✍️</strong></a><strong> 2021). </strong><br>There’s no object, as far as we, humans, know, that is faster than the <a href="https://en.wikipedia.org/wiki/Speed_of_light">speed of light</a>. In this <a href="http://blog.cloudflare.com/fastest-internet/"><strong>blog post</strong></a> you’ll get a sense of the physical limits of Internet speeds (“the speed of light is really slow”). How it all works through electrons through wires, lasers blasting data down fiber optic cables, and how building a waitless Internet is hard. <br>We go on to explain the factors that go into building our fast global network: bandwidth, latency, reliability, caching, cryptography, DNS, preloading, cold starts, and more; and how Cloudflare zeroes in on the most powerful number there is: zero. And here’s a challenge, there are a few movies, books, board game references hidden in the <a href="http://blog.cloudflare.com/fastest-internet/">post</a> for you to find.<br><br></li></ul><blockquote><em>“People ask me to predict the future, when all I want to do is prevent it. Better yet, build it. Predicting the future is much too easy, anyway. You look at the people around you, the street you stand on, the visible air you breathe, and predict more of the same. To hell with more. I want better.”</em><br>— <strong>Ray Bradbury</strong>, from Beyond 1984: The People Machines</blockquote><p></p><ul><li><strong>Securing the post-quantum world (</strong><a href="http://blog.cloudflare.com/securing-the-post-quantum-world/"><strong>✍️</strong></a><strong> 2020).</strong><br>This one is more about the future of the Internet. We have many post-quantum <a href="http://blog.cloudflare.com/tag/post-quantum/">related posts</a>, including the recent standardization one (‘<a href="http://blog.cloudflare.com/nist-post-quantum-surprise/">NIST’s pleasant post-quantum surprise</a>’), but <a href="http://blog.cloudflare.com/securing-the-post-quantum-world/"><strong>here</strong></a> you have an easy-to-understand explanation of a complex but crucial for the future of the Internet topic. More on those challenges and opportunities in 2022 <a href="http://blog.cloudflare.com/post-quantum-future/">here</a>. <br>The sum up is: “Quantum computers are coming that will have the ability to break the cryptographic mechanisms we rely on to secure modern communications, but there is hope”. For a quantum computing starting point, check: <a href="http://blog.cloudflare.com/the-quantum-menace/">The Quantum Menace</a>.<br><br></li><li><strong>SAD DNS Explained (</strong><a href="http://blog.cloudflare.com/sad-dns-explained/"><strong>✍️</strong></a><strong> 2020).</strong> <br>A 2020 attack against the Domain Name System (DNS) called SAD DNS (Side channel AttackeD DNS) leveraged features of the networking stack in modern operating systems. It’s a good excuse to <a href="http://blog.cloudflare.com/sad-dns-explained/"><strong>explain</strong></a> how the DNS protocol and spoofing <a href="http://blog.cloudflare.com/dns-encryption-explained/">work</a>, and how the industry can prevent it — another post expands on improving <a href="http://blog.cloudflare.com/oblivious-dns/">DNS privacy</a> with Oblivious DoH in 1.1.1.1.<br><br></li><li><strong>Privacy needs to be built into the Internet (</strong><a href="http://blog.cloudflare.com/internet-privacy/"><strong>✍️</strong></a><strong> 2020)</strong><br>A bit of history is always interesting and of value (at least for me). To launch one of our <a href="http://blog.cloudflare.com/tag/privacy/">Privacy</a> Weeks, in 2020, here’s a <a href="http://blog.cloudflare.com/internet-privacy/"><strong>general view</strong></a> to the three different phases of the Internet. Until the 1990s the race was for connectivity. With the introduction of SSL in 1994, the Internet moved to a second phase where security became paramount (it helped create the dotcom rush and the secure, online world we live in today). Now, it’s all about the Phase 3 of the Internet we’re helping to build: always on, always secure, always private.<br><br></li><li><strong>50 Years of The Internet. Work in Progress to a Better Internet (</strong><a href="http://blog.cloudflare.com/50-years-of-the-internet-work-in-progress-to-a-better-internet/"><strong>✍️</strong></a> <strong>2019)</strong><br>In 2019, we were <a href="http://blog.cloudflare.com/50-years-of-the-internet-work-in-progress-to-a-better-internet/"><strong>celebrating 50</strong></a> years from when the very first network packet took flight from the Los Angeles campus at UCLA to the Stanford Research Institute (SRI) building in Palo Alto. Those two California sites had kicked-off the world of packet networking, on the ARPANET, and of the modern Internet as we use and know it today. Here we go through some Internet history. <br>This reminds me of this December 2021 <a href="https://twitter.com/Cloudflare/status/1471918044343676936">conversation</a> about how the Web began, 30 years earlier. Cloudflare CTO John Graham-Cumming meets Dr. Ben Segal, early Internet pioneer and CERN's first official TCP/IP Coordinator, and Francois Fluckiger, director of the CERN School of Computing. <a href="https://twitter.com/Cloudflare/status/1471918044343676936">Here</a>, we learn how the World Wide Web became an open source project.<br><br></li><li><strong>Welcome to Crypto Week (</strong><a href="http://blog.cloudflare.com/crypto-week-2018/"><strong>✍️</strong></a> <strong>2018).</strong><br>If you want to know why cryptography is so important for the Internet, <a href="http://blog.cloudflare.com/crypto-week-2018/"><strong>here’s</strong></a> a good place to start. The Internet, with all of its marvels in connecting people and ideas, needs an upgrade, and one of the tools that can make things better is cryptography. There’s also a more mathematical <a href="http://blog.cloudflare.com/privacy-pass-the-math/">privacy pass protocol</a> related perspective (that is the basis of the work to eliminate <a href="http://blog.cloudflare.com/eliminating-captchas-on-iphones-and-macs-using-new-standard/">CAPTCHAs</a>).<br><br></li><li><strong>Why TLS 1.3 isn't in browsers yet (</strong><a href="http://blog.cloudflare.com/why-tls-1-3-isnt-in-browsers-yet/"><strong>✍️</strong></a><strong> 2017).</strong><br>It’s all <a href="http://blog.cloudflare.com/why-tls-1-3-isnt-in-browsers-yet/"><strong>about</strong></a>: “Upgrading a security protocol in an ecosystem as complex as the Internet is difficult. You need to update clients and servers and make sure everything in between continues to work correctly. The Internet is in the middle of such an upgrade right now.” More on that from 2021 <a href="http://blog.cloudflare.com/handshake-encryption-endgame-an-ech-update/">here</a>: <em>Handshake Encryption: Endgame (an ECH update)</em>.<br><br></li><li><strong>How to build your own public key infrastructure (</strong><a href="http://blog.cloudflare.com/how-to-build-your-own-public-key-infrastructure/"><strong>✍️</strong></a> <strong>2015).</strong><br>A way of <a href="http://blog.cloudflare.com/how-to-build-your-own-public-key-infrastructure/"><strong>getting to know</strong></a> how a major part of securing a network as geographically diverse as Cloudflare’s is protecting data as it travels between datacenters. “Great security architecture requires a defense system with multiple layers of protection”. From the same year, <a href="http://blog.cloudflare.com/why-its-harder-to-forge-a-sha-1-certificate-than-it-is-to-find-a-sha-1-collision/">here’s</a> something about digital signatures being the bedrock of trust.<br><br></li><li><strong>A (Relatively Easy To Understand) Primer on Elliptic Curve Cryptography (</strong><a href="http://blog.cloudflare.com/a-relatively-easy-to-understand-primer-on-elliptic-curve-cryptography/"><strong>✍️</strong></a><strong> 2013).</strong><br>Also thinking of how the Internet will continue to work for years to come, <a href="http://blog.cloudflare.com/a-relatively-easy-to-understand-primer-on-elliptic-curve-cryptography/"><strong>here’s</strong></a> a very complex topic made simple about one of the most powerful but least understood types of cryptography in wide use.<br><br></li><li><strong><strong><strong>Why Google Went Offline Today and a Bit about How the Internet Works (</strong><a href="http://blog.cloudflare.com/why-google-went-offline-today-and-a-bit-about/"><strong>✍️</strong></a><strong> 2012).</strong></strong></strong><br>We had several similar blog posts over the years, but <a href="http://blog.cloudflare.com/why-google-went-offline-today-and-a-bit-about/">this 10-year old one</a> from <a href="https://cloudflare.tv/event/0o813MtDM4EtrZXttzuJO">Tom Paseka</a> set the tone on how we could give a good technical explanation for something that was impacting so many. Here ​​Internet routing, route leakages are discussed and it all ends on a relevant note: “Just another day in our ongoing efforts to #savetheweb.” Quoting from someone in the company for nine years: “This blog was the one that first got me interested in Cloudflare”.</li></ul><p>Again, if you like Twitter threads, <a href="https://twitter.com/grittygrease/status/1555201358357270528"><strong>this</strong></a> recent Nick Sullivan one starts with an announcement (Cloudflare now allows <a href="http://blog.cloudflare.com/experiment-with-pq/">experiments with post-quantum cryptography</a>) and goes on explaining what some of the more relevant Internet acronyms mean. Example: TLS, or Transport Layer Security, it’s the ubiquitous encryption and authentication protocol that protects web requests online.</p><h3 id="blast-from-the-past-some-history-">Blast from the past (some history)</h3><p>A few also recently referenced blog posts from the past, some more technical than others.</p><ul><li><strong>Introducing DNS Resolver, 1.1.1.1 (not a joke) (</strong><a href="http://blog.cloudflare.com/dns-resolver-1-1-1-1/"><strong>✍️</strong></a><strong> 2018).</strong><br>The first consumer-focused service Cloudflare has ever released, our DNS resolver, <a href="https://1.1.1.1/">1.1.1.1</a> — a recursive DNS service — was <a href="http://blog.cloudflare.com/announcing-1111/">launched</a> on April 1, 2018, and <a href="http://blog.cloudflare.com/dns-resolver-1-1-1-1/"><strong>this is</strong></a> the technical explanation. With this offering, we started fixing the foundation of the Internet by building a faster, more secure and privacy-centric public DNS resolver. And, just this month, we’ve added <a href="http://blog.cloudflare.com/geoexit-improving-warp-user-experience-larger-network/">privacy proofed features</a> (a geolocation accuracy “pizza test” included). <br><br></li><li><strong>Cloudflare goes InterPlanetary - Introducing Cloudflare’s IPFS Gateway (</strong><a href="http://blog.cloudflare.com/distributed-web-gateway/"><strong>✍️</strong></a><strong> 2018).</strong><br>We <a href="http://blog.cloudflare.com/distributed-web-gateway/"><strong>introduced</strong></a> Cloudflare’s IPFS Gateway, an easy way to access content from the InterPlanetary File System (IPFS). This served as the platform for many new, at the time, highly-reliable and security-enhanced web applications. It was the first product to be released as part of our <a href="https://www.cloudflare.com/distributed-web-gateway">Distributed Web Gateway</a> project and is a different perspective from the traditional web. <br>IPFS is a peer-to-peer file system composed of thousands of computers around the world, each of which stores files on behalf of the network. And, yes, it can be used as a method for a possible Mars (Moon, etc.) Internet in the future. About that, the same goes for code that will need to be running on Mars, something we mention about Workers <a href="http://blog.cloudflare.com/cloudflare-workers-unleashed/">here</a>.<br><br></li><li><strong>LavaRand in Production: The Nitty-Gritty Technical Details (</strong><a href="http://blog.cloudflare.com/lavarand-in-production-the-nitty-gritty-technical-details"><strong>✍️</strong></a><strong> 2017).</strong><br>Our lava lamps wall in the San Francisco office is much more than <a href="https://www.fastcodesign.com/90137157/the-hardest-working-office-design-in-america-encrypts-your-data-with-lava-lamps">a wall of lava lamps</a> (the YouTuber Tom Scott did a 2017 <a href="https://www.youtube.com/watch?v=1cUUfMeOijg">video</a> about it) and in <a href="http://blog.cloudflare.com/lavarand-in-production-the-nitty-gritty-technical-details"><strong>this blog</strong></a> we explain the in-depth look at the technical details (there’s a less technical <a href="http://blog.cloudflare.com/randomness-101-lavarand-in-production/">one</a> on how randomness in cryptography works). <br><br></li><li><strong><strong><strong>Introducing Cloudflare Workers (</strong><a href="http://blog.cloudflare.com/introducing-cloudflare-workers/"><strong>✍️</strong></a><strong> 2017).</strong></strong></strong><br>There are several announcements each year, but <a href="http://blog.cloudflare.com/introducing-cloudflare-workers/">this blog</a> (associated with the explanation, <a href="http://blog.cloudflare.com/code-everywhere-cloudflare-workers/">Code Everywhere: Why We Built Cloudflare Workers</a>) was referenced this week by some as one of those with a clear impact. It was when we started making <a href="https://www.cloudflare.com/network/">Cloudflare's network</a> programmable. In 2018, <a href="http://blog.cloudflare.com/cloudflare-workers-unleashed/">Workers</a> was available to everyone and, in 2019, we registered the trademark for <a href="http://blog.cloudflare.com/the-network-is-the-computer/">The Network is the Computer®</a>, to encompass how Cloudflare is using its network to pave the way for the future of the Internet.<br><br></li><li><strong>What's the story behind the names of CloudFlare's name servers? (</strong><a href="http://blog.cloudflare.com/whats-the-story-behind-the-names-of-cloudflares-name-servers/"><strong>✍️</strong></a><strong> 2013)</strong><br>Another one referenced this week is the answer to the question we got often back in 2013: what the names of our nameservers mean. <a href="http://blog.cloudflare.com/whats-the-story-behind-the-names-of-cloudflares-name-servers/"><strong>Here's the story</strong></a> — there’s even an Apple co-founder <a href="https://cloudflare.tv/event/4Zvur2v6gLs8lg4Zzs7iX7">Steve Wozniak</a> tribute.</li></ul>]]></content:encoded></item><item><title><![CDATA[Cloudflare Support Portal gets an overhaul]]></title><description><![CDATA[The Cloudflare Support team has launched a new Support Portal. The portal will give you access to self-help resources, diagnostics with troubleshooting guides, and will provide for easier ticket submission]]></description><link>https://blog.cloudflare.com/cloudflare-support-portal-gets-an-overhaul/</link><guid isPermaLink="false">62fa2000eae9c6000af21ba6</guid><category><![CDATA[Customer Support]]></category><category><![CDATA[Support Portal]]></category><category><![CDATA[Support]]></category><dc:creator><![CDATA[Meghan Bevill]]></dc:creator><pubDate>Tue, 16 Aug 2022 13:00:00 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/unnamed-5.png" medium="image"/><content:encoded><![CDATA[<figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/unnamed-4.png" class="kg-image" alt="Cloudflare Support Portal gets an overhaul"></figure><img src="http://blog.cloudflare.com/content/images/2022/08/unnamed-5.png" alt="Cloudflare Support Portal gets an overhaul"><p>The Cloudflare Support team is excited to announce the launch of our brand-new Customer Support Portal. When our customers open support tickets, we understand that they want quick and accurate responses from us. For those of you who have opened a support ticket in the past, we are certain you will notice the improvements we've made! The new Support Portal lives where our ticket submission form has always been, <a href="https://dash.cloudflare.com/?to=/:account/support">dash.cloudflare.com/support</a>, but that's where the similarities between the old and the new one end.</p><h3 id="what-can-you-expect-in-the-new-portal">What can you expect in the new portal?</h3><p>The new Support Portal will help you solve your problems quickly and effectively, by getting you on the fastest path to resolution. In some cases, the most efficient way to resolve your issue will be to use our self-help resources or our machine learning-trained Support Bot. Other times, the most efficient way to resolve your issue will be by working with one of our Support Engineers via ticket, phone or chat, depending on your plan type. Regardless of how we help you solve your issue, we will have more context about the products you are using and your issue up front, reducing time-consuming back and forth.</p><p>The new portal has several features that will make it easier for you to access the support you need, including:</p><ul><li>Fast and secure ticket submission for verified Cloudflare users</li><li>An easier-to-use interface that serves relevant resources based on your issue summary</li><li>Machine learning-powered Support Bot to run diagnostics and serve targeted help guides</li></ul><p>Everyone is encouraged to begin using our new portal. Tickets submitted through our legacy form are typically solved faster than tickets emailed to us, and we expect the updates in our new form to help us resolve your issues even faster!</p><p>If you are ready to be one of the first people to take advantage of our new Support Portal, you can now opt in and begin using the new experience to access resources and submit tickets. Just hit the Support dropdown in your <a href="https://dash.cloudflare.com/?account=support">dashboard</a> and click Contact Support.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image3-7.png" class="kg-image" alt="Cloudflare Support Portal gets an overhaul"></figure><p>Below is a preview of what you can expect with the new experience.</p><h3 id="relevant-self-help-resources-at-your-fingertips">Relevant self-help resources at your fingertips</h3><p>The biggest change you’ll notice from our old ticket submission form is that we've made it easier to get help. First, we link you directly to relevant resources and the ticket submission form immediately upon clicking “Contact Support”. You no longer have to navigate through multiple steps to get your problem resolved. Second, we’ve moved to a full-page experience allowing us to curate a selection of support articles and help guides targeting your specific problem, making it easier for you to find answers to your questions. Of course, there will still be times when you need to submit a support ticket, but if we have resources that address your problem, we want you to be able to find that information easily.</p><p>All the details you provide when searching for articles in the portal will be captured and added to your ticket if you are not able to find the answers to your questions.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image1-16.png" class="kg-image" alt="Cloudflare Support Portal gets an overhaul"></figure><h3 id="take-advantage-of-our-support-bot">Take advantage of our Support Bot</h3><p>Our machine learning-powered Support Bot has been integrated into the new portal to deliver a customized experience that identifies your specific problem. Support Bot has been helping our Support Engineers work more efficiently for years, and now we’re making some of this functionality customer-facing so that you can benefit from these efficiencies as well.</p><p>Within the portal, the Support Bot will run diagnostics (if your issue is domain-related), assess the issue summary you entered, and provide you with help guides to address the root cause of your problem. The more information you are able to provide, the better our bot can direct you to the resources most pertinent to your issue. This gives you the chance to solve your issue on the spot, rather than waiting for a response to your ticket.</p><p>For each issue submitted through the portal, our Support Bot can perform one of two actions. If your issue is domain-specific, the bot will run a set of diagnostics against your domain that check for common configuration issues. If any issue is detected, the bot will display the issue and a suggested solution. Regardless of whether your issue is domain-specific, the bot will also analyze the issue summary you’ve entered against our ensemble of Natural Language Processing models and keyword searches. The bot is trained on thousands of historic customer tickets to differentiate between specific customer issues. We retrain the model on a regular basis to ensure it is consistently learning from new and emerging issues. If the bot detects keywords in your summary that map to a relevant issue, it will present a known solution for that issue.</p><p>The solutions the bot surfaces are based on how successfully these resources resolved issues previously, and we will continue to refine the bot’s responses and solutions based on a couple of key success metrics. We consider a recommendation successful if a customer doesn’t need to ultimately open a ticket or if they acknowledge that a resource was helpful by voting on the page. We will evaluate this data along with any information you provide on why specific content wasn’t helpful, and make iterative improvements to the bot every time we retrain it.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image2-10.png" class="kg-image" alt="Cloudflare Support Portal gets an overhaul"></figure><h3 id="fast-and-secure-ticket-submission">Fast and secure ticket submission</h3><p>While we have a ton of helpful content for a wide range of problems, we know there will be instances where you need to speak to one of our very experienced Support Engineers. For plan types that include ticket support, we have built our ticket submission flow into the portal and introduced new features to make the experience more efficient. The first step for our Support Engineers in resolving most issues is for us to verify the identity of account users and admins. The new process ensures that tickets are only submitted by verified account users and admins, reducing some back and forth and allowing us to start working on your issue right away.</p><p>Along with this verification step, the new portal will collect detailed information about your problem up front, including issue category and impact level. These details will help route your ticket to the Support Engineer most knowledgeable in the area of your issue and enable that engineer to begin work on your ticket more quickly and without having to come to you with additional questions.</p><h3 id="how-to-try-the-new-experience">How to try the new experience</h3><p>To take advantage of these improvements, we encourage everyone to use the new Support Portal as the starting point for troubleshooting your issues.</p><p>Over the next few months, we will be rolling out the new portal to all plan types, starting with an opt-in period where you can pilot the new experience. Once we are satisfied the portal is working as intended, we will close the opt-in phase and release the portal to all customers. At that point, we will begin redirecting emails received at our main support email addresses (support at cloudflare.com and billing at cloudflare.com) to the Support Portal so that they can be triaged, and resolved quicker and more efficiently. We are excited to start implementing these changes and are confident that these steps are the first of many planned in making your support experience as efficient and effective as possible. We can’t wait for you to check it out!</p><p>To start using the new portal today, you can opt in from your <a href="https://dash.cloudflare.com/?account=support">dashboard</a>. Let us know what you think with the <a href="https://forms.gle/DdP6xdbeyfUDd7QL8">feedback form</a> included at the top of the new portal.</p>]]></content:encoded></item><item><title><![CDATA[Crawler Hints supports Microsoft’s IndexNow in helping users find new content]]></title><description><![CDATA[Cloudflare is uniquely positioned to help give crawlers hints about when they should recrawl, if new content has been added, or if content on a site has recently changed]]></description><link>https://blog.cloudflare.com/crawler-hints-supports-microsofts-indexnow-in-helping-users-find-new-content/</link><guid isPermaLink="false">62f5e702de861a000a8df0bb</guid><category><![CDATA[Crawler Hints]]></category><category><![CDATA[Search Engine]]></category><category><![CDATA[Bots]]></category><category><![CDATA[Cache]]></category><category><![CDATA[Product News]]></category><dc:creator><![CDATA[Alex Krivit]]></dc:creator><pubDate>Fri, 12 Aug 2022 16:30:20 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/image2-8.png" medium="image"/><content:encoded><![CDATA[<figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image2-9.png" class="kg-image" alt="Crawler Hints supports Microsoft’s IndexNow in helping users find new content"></figure><img src="http://blog.cloudflare.com/content/images/2022/08/image2-8.png" alt="Crawler Hints supports Microsoft’s IndexNow in helping users find new content"><p>The web is constantly changing. Whether it’s news or updates to your social feed, it’s a constant flow of information. As a user, that’s great. But have you ever stopped to think how search engines deal with all the change?</p><p>It turns out, they “index” the web on a regular basis — sending bots out, to constantly crawl webpages, looking for changes. Today, bot traffic accounts for about <a href="https://radar.cloudflare.com/?date_filter=last_30_days">30% of total traffic</a> on the Internet, and given how foundational search is to using the Internet, it should come as no surprise that search engine bots make up a large proportion of that what might come as a surprise is how inefficient the model is, though: we estimate that over <a href="http://blog.cloudflare.com/crawler-hints-how-cloudflare-is-reducing-the-environmental-impact-of-web-searches/">50% of crawler traffic is wasted effort</a>.</p><p>This has a huge impact. There’s all the additional capacity that owners of websites need to bake into their site to absorb the bots crawling all over it. There’s the transmission of the data. There’s the CPU cost of running the bots. And when you’re running at the scale of the Internet, all of this has a pretty big environmental footprint.</p><p>Part of the problem, though, is nobody had really stopped to ask: maybe there’s a better way?</p><p>Right now, the model for indexing websites is the same as it has been since the 1990s: a “pull” model, where the search engine sends a crawler out to a website after a predetermined amount of time. During Impact Week last year, we asked: what about flipping the model on its head? What about moving to a push model, where a website could simply ping a search engine to let it know an update had been made?</p><p>There are a heap of advantages to such a model. The website wins: it’s not dealing with unnecessary crawls. It also makes sure that as soon as there’s an update to its content, it’s reflected in the search engine — it doesn’t need to wait for the next crawl. The website owner wins because they don't need to manage distinct search engine crawl submissions. The search engine wins, too: it saves money on crawl costs, and it can make sure it gets the latest content.</p><p>Of course, this needs work to be done on both sides of the equation. The websites need a mechanism to alert the search engines; and the search engines need a mechanism to receive the alert, so they know when to do the crawl.</p><h3 id="crawler-hints-cloudflare-s-solution-for-websites">Crawler Hints — Cloudflare’s Solution for Websites</h3><p>Solving this problem is why we <a href="http://blog.cloudflare.com/crawler-hints-how-cloudflare-is-reducing-the-environmental-impact-of-web-searches/">launched Crawler Hints</a>. Cloudflare sits in a unique position on the Internet — we’re serving on average 36 million HTTP requests per second. That represents <em>a lot of websites</em>. It also means we’re uniquely positioned to help solve this problem:  to help give crawlers hints about when they should recrawl if new content has been added or if content on a site has recently changed.</p><p>With Crawler Hints, we send signals to web indexers based on cache data and origin status codes to help them understand when content has likely changed or been added to a site. The aim is to increase the number of relevant crawls as well as drastically reduce the number of crawls that don’t find fresh content, saving bandwidth and compute for both indexers and sites alike, and improving the experience of using the search engines.</p><p>But, of course, that’s just half the equation.</p><h3 id="indexnow-protocol-the-search-engine-moves-from-pull-to-push">IndexNow Protocol — the Search Engine Moves from Pull to Push</h3><p>Websites alerting the search engine about changes is useless if the search engines aren’t listening — and they simply continue to crawl the way they always have. Of course, search engines are incredibly complicated, and changing the way they operate is no easy task.</p><p>The IndexNow Protocol is a standard developed by Microsoft, Seznam.cz and Yandex, and it represents a major shift in the way search engines operate. Using IndexNow, search engines have a mechanism by which they can receive signals from Crawler Hints. Once they have that signal, they can shift their crawlers from a pull model to a push model.</p><p>In a recent update, <a href="https://blogs.bing.com/webmaster/august-2022/IndexNow-adoption-gains-momentum">Microsoft has announced</a> that millions of websites are now using IndexNow to signal to search engine crawlers when their content needs to be crawled and IndexNow was used to<strong> index/crawl about 7% of all new URLs</strong> <strong>clicked</strong> when someone is selecting from web search results.</p><p>On the Cloudflare side, since the release of Crawler Hints in October 2021, Crawler Hints has processed about <strong>six-hundred-billion</strong> signals to IndexNow.</p><p>That’s a lot of saved crawls.</p><h3 id="how-to-enable-crawler-hints">How to enable Crawler Hints</h3><p>By enabling Crawler Hints on your website, with the simple click of a button, Cloudflare will take care of signaling to these search engines when your content has changed via the <a href="https://www.indexnow.org/">IndexNow</a> API. You don’t need to do anything else!</p><p>Crawler Hints is free to use and available to all Cloudflare customers. If you’d like to see how Crawler Hints can benefit how your website is indexed by the world's biggest search engines, please feel free to opt-into the service by:</p><ol><li>Sign in to your Cloudflare Account.</li><li>In the dashboard, navigate to the Cache tab.</li><li>Click on the Configuration section.</li><li>Locate the Crawler Hints and enable.</li></ol><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image1-15.png" class="kg-image" alt="Crawler Hints supports Microsoft’s IndexNow in helping users find new content"></figure><p>Upon enabling Crawler Hints, Cloudflare will share when content on your site has changed and needs to be re-crawled with search engines using the IndexNow protocol (<a href="http://blog.cloudflare.com/from-0-to-20-billion-how-we-built-crawler-hints/">this blog</a> can help if you’re interested in finding out more about how the mechanism works).</p><h3 id="what-s-next">What’s Next?</h3><p>Going forward, because the benefits are so substantial for site owners, search operators, and the environment, we plan to start defaulting Crawler Hints on for all our customers. We’re also hopeful that Google, the world’s largest search engine and most wasteful user of Internet resources, will adopt IndexNow or a similar standard and lower the burden of search crawling on the planet.</p><p>When we think of helping to build a better Internet, this is exactly what comes to mind: creating and supporting standards that make it operate better, greener, faster. We’re really excited about the work to date, and will continue to work to improve the signaling to ensure the most valuable information is being sent to the search engines in a timely manner. This includes incorporating additional signals such as etags, last-modified headers, and content hash differences. Adding these signals will help further inform crawlers when they should reindex sites, and how often they need to return to a particular site to check if it’s been changed. This is only the beginning. We will continue testing more signals and working with industry partners so that we can help crawlers run efficiently with these hints.</p><p>And finally: if you’re on Cloudflare, and you’d like to be part of this revolution in how search engines operate on the web (it’s free!), simply follow the instructions in the section above.</p>]]></content:encoded></item><item><title><![CDATA[2022 attacks! An August reading list to go “Shields Up”]]></title><description><![CDATA[In 2022, cybersecurity, more than ever, is a must-have for those who don’t want to take chances on getting caught in a cyberattack with difficult to deal with consequences. Here’s a reading list what you need to know about attacks that is also a guide on how to be protected]]></description><link>https://blog.cloudflare.com/2022-attacks-an-august-reading-list-to-go-shields-up/</link><guid isPermaLink="false">62f34ecade861a000a8dee39</guid><category><![CDATA[Reading List]]></category><category><![CDATA[Security]]></category><category><![CDATA[Attacks]]></category><category><![CDATA[DDoS]]></category><category><![CDATA[Ransom Attack]]></category><category><![CDATA[Phishing]]></category><dc:creator><![CDATA[João Tomé]]></dc:creator><pubDate>Thu, 11 Aug 2022 13:00:00 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/image4-1.png" medium="image"/><content:encoded><![CDATA[<!--kg-card-begin: markdown--><img src="http://blog.cloudflare.com/content/images/2022/08/image4-1.png" alt="2022 attacks! An August reading list to go “Shields Up”"><p><em><small>This post is also available in <a href="http://blog.cloudflare.com/zh-cn/2022-attacks-an-august-reading-list-to-go-shields-up-zh-cn/">简体中文</a>, <a href="http://blog.cloudflare.com/ja-jp/2022-attacks-an-august-reading-list-to-go-shields-up-ja-jp/">日本語</a>, <a href="http://blog.cloudflare.com/de-de/2022-attacks-an-august-reading-list-to-go-shields-up-de-de/">Deutsch</a>, <a href="http://blog.cloudflare.com/fr-fr/2022-attacks-an-august-reading-list-to-go-shields-up-fr-fr/">Français</a>, and <a href="http://blog.cloudflare.com/2022-attacks-an-august-reading-list-to-go-shields-up-es-es/">Español</a>.</small></em></p>
<!--kg-card-end: markdown--><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image4-2.png" class="kg-image" alt="2022 attacks! An August reading list to go “Shields Up”"></figure><p>In 2022, cybersecurity is a must-have for those who don’t want to take chances on getting caught in a cyberattack with difficult to deal consequences. And with a war in Europe (<a href="http://blog.cloudflare.com/tag/ukraine/">Ukraine</a>) still going on, cyberwar also doesn’t show signs of stopping in a time when there never were so many people online, 4.95 billion in early 2022, 62.5% of the world’s total population (<a href="https://datareportal.com/reports/digital-2022-global-overview-report">estimates</a> say it grew around 4% during 2021 and <a href="https://datareportal.com/reports/digital-2021-global-overview-report">7.3%</a> in 2020).</p><p>Throughout the year we, at Cloudflare, have been making new announcements of products, solutions and initiatives that highlight the way we have been preventing, mitigating and constantly learning, over the years, with several thousands of small and big cyberattacks. Right now, we block an average of 124 billion cyber threats per day. The more we deal with attacks, the more we know how to stop them, and the easier it gets to find and deal with new threats — and for customers to forget we’re there, protecting them.</p><p>In 2022, we have been onboarding many customers while they’re being attacked, something we know well from the past (<a href="https://www.cloudflare.com/case-studies/wikimedia-foundation/">Wikimedia/Wikipedia</a> or <a href="https://www.cloudflare.com/case-studies/eurovision/">Eurovision</a> are just two case-studies of <a href="https://www.cloudflare.com/case-studies">many</a>, and last year there was a Fortune Global 500 company example we <a href="http://blog.cloudflare.com/ransom-ddos-attacks-target-a-fortune-global-500-company/">wrote about</a>). Recently, we dealt and did a <a href="http://blog.cloudflare.com/2022-07-sms-phishing-attacks/">rundown</a> about an SMS phishing attack.</p><p>Providing services for <a href="https://w3techs.com/technologies/overview/proxy/all">almost 20%</a> of websites online and to millions of Internet properties and customers using our global network in more than <a href="http://blog.cloudflare.com/new-cities-april-2022-edition/">270 cities</a> (recently we arrived to <a href="http://blog.cloudflare.com/cloudflare-deployment-in-guam/">Guam</a>) also plays a big role. For example, in Q1’22 Cloudflare blocked an average of 117 billion cyber threats each day (much more than in previous quarters).</p><p>Now that August is here, and many in the Northern Hemisphere are enjoying the summer and vacations, let’s do a reading list that is also a sum up focused on cyberattacks that also gives, by itself, some 2022 guide on this more than ever relevant area.</p><h2 id="war-cyberwar-attacks-increasing">War &amp; Cyberwar: Attacks increasing</h2><p>But first, some context. There are all sorts of attacks, but they have been generally speaking increasing and just to give some of our data regarding <a href="http://blog.cloudflare.com/ddos-attack-trends-for-2022-q2/">DDoS attacks in 2022 Q2</a>: ​​application-layer attacks increased by 72% YoY (Year over Year) and network-layer DDoS attacks increased by 109% YoY.</p><p>The US government gave “warnings” back in March, after the war in Ukraine started, to all in the country but also allies and partners to be aware of the need to “enhance cybersecurity”. The US Cybersecurity and Infrastructure Security Agency (CISA) created the <a href="https://www.cisa.gov/shields-up">Shields Up</a> initiative, given how the “Russia’s invasion of Ukraine could impact organizations both within and beyond the region”. The <a href="http://blog.cloudflare.com/shields-up-free-cloudflare-services-to-improve-your-cyber-readiness/#:~:text=National%20Cyber%20Security%20Center">UK</a> and <a href="https://www.meti.go.jp/press/2021/02/20220221003/20220221003.html">Japan</a>, among others, also issued warnings.</p><p>That said, here are the two first and more general about attacks reading list suggestions:</p><p><strong>Shields up: free Cloudflare services to improve your cyber readiness (</strong><a href="http://blog.cloudflare.com/shields-up-free-cloudflare-services-to-improve-your-cyber-readiness/"><strong>✍️</strong></a><strong>)</strong><br>After the war started and governments released warnings, we did this free Cloudflare services cyber readiness sum up <a href="http://blog.cloudflare.com/shields-up-free-cloudflare-services-to-improve-your-cyber-readiness/">blog post</a>. If you’re a seasoned IT professional or a novice website operator, you can see a variety of services for websites, apps, or APIs, including DDoS mitigation and protection of teams or even personal devices (from phones to routers). If this resonates with you, this announcement of collaboration to simplify the adoption of Zero Trust for IT and security teams could also be useful: <a href="http://blog.cloudflare.com/cloudflare-crowdstrike-partnership/">CrowdStrike’s endpoint security meets Cloudflare’s Zero Trust Services</a>.</p><p><strong>In Ukraine and beyond, what it takes to keep vulnerable groups online (</strong><a href="http://blog.cloudflare.com/in-ukraine-and-beyond-what-it-takes-to-keep-vulnerable-groups-online/"><strong>✍️</strong></a><strong>)</strong><br>This <a href="http://blog.cloudflare.com/in-ukraine-and-beyond-what-it-takes-to-keep-vulnerable-groups-online/">blog post</a> is focused on the eighth anniversary of our <a href="https://www.cloudflare.com/galileo/">Project Galileo</a>, that has been helping human-rights, journalism and non-profits public interest organizations or groups. We highlight the trends of the past year, including the dozens of organizations related to <a href="http://blog.cloudflare.com/tag/Ukraine">Ukraine</a> that were onboarded (many while being attacked) since the war started. Between July 2021 and May 2022, we’ve blocked an average of nearly 57.9 million cyberattacks per day, an increase of nearly 10% over last year in a total of 18 billion attacks.</p><p>In terms of attack methods to Galileo protected organizations, the largest fraction (28%) of mitigated requests were classified as “HTTP Anomaly”, with 20% of mitigated requests tagged as <a href="https://www.cloudflare.com/learning/security/threats/sql-injection/">SQL injection or SQLi attempts</a> (to target databases) and nearly 13% as attempts to exploit specific <a href="https://www.cve.org/">CVEs</a> (publicly disclosed cybersecurity vulnerabilities) — you can find more insights about those <a href="http://blog.cloudflare.com/tag/cve/">here</a>, including the <a href="http://blog.cloudflare.com/waf-mitigations-spring4shell/">Spring4Shell</a> vulnerability, the <a href="http://blog.cloudflare.com/tag/log4j/">Log4j</a> or the <a href="http://blog.cloudflare.com/cloudflare-customers-are-protected-from-the-atlassian-confluence-cve-2022-26134/">Atlassian</a> one.</p><p>And now, without further ado, here’s the full reading list/attacks guide where we highlight some blog posts around four main topics:</p><h2 id="1-ddos-attacks-solutions">1. DDoS attacks &amp; solutions </h2><figure class="kg-card kg-image-card kg-card-hascaption"><img src="http://blog.cloudflare.com/content/images/2022/08/image5-2.png" class="kg-image" alt="2022 attacks! An August reading list to go “Shields Up”"><figcaption>The most powerful botnet to date, <a href="http://blog.cloudflare.com/mantis-botnet/">Mantis</a>.</figcaption></figure><p><strong>Cloudflare mitigates 26 million request per second DDoS attack (</strong><a href="http://blog.cloudflare.com/26m-rps-ddos/"><strong>✍️</strong></a><strong>)</strong><br><a href="https://www.cloudflare.com/en-gb/learning/ddos/what-is-a-ddos-attack/">Distributed Denial of Service (DDoS)</a> are the bread and butter of <a href="https://portswigger.net/daily-swig/nation-state-threat-how-ddos-over-tcp-technique-could-amplify-attacks">state-based</a> attacks, and we’ve been <a href="http://blog.cloudflare.com/deep-dive-cloudflare-autonomous-edge-ddos-protection/">automatically</a> detecting and mitigating them. Regardless of which country initiates them, bots are all around the world and <a href="http://blog.cloudflare.com/26m-rps-ddos/">in this blog post</a> you can see a specific example on how big those attacks can be (in this case the attack targeted a customer website using Cloudflare’s Free plan). We’ve named this most powerful botnet to date, <a href="http://blog.cloudflare.com/mantis-botnet/">Mantis</a>.</p><p>That said, we also explain that although most of the attacks are small, e.g. cyber vandalism, even small attacks can severely impact unprotected Internet properties.</p><p><strong>DDoS attack trends for 2022 Q2 (</strong><a href="http://blog.cloudflare.com/ddos-attack-trends-for-2022-q2/"><strong>✍️</strong></a><strong>)</strong><br>We already mentioned how application (72%) and network-layer (109%) attacks have been growing year over year — in the latter, attacks of 100 Gbps and larger increased by 8% QoQ, and attacks lasting more than 3 hours increased by 12% QoQ. <a href="http://blog.cloudflare.com/ddos-attack-trends-for-2022-q2/"><strong>Here</strong></a> you can also find interesting trends, like how Broadcast Media companies in Ukraine were the most targeted in Q2 2022 by DDoS attacks. In fact, all the top five most attacked industries are all in online/Internet media, publishing, and broadcasting.</p><p><strong>Cloudflare customers on Free plans can now also get real-time DDoS alerts</strong><a href="http://blog.cloudflare.com/free-ddos-alerts/"><strong> </strong></a><strong>(</strong><a href="http://blog.cloudflare.com/free-ddos-alerts/"><strong>✍️</strong></a><strong>)</strong><br>A DDoS is cyber-attack that attempts to disrupt your online business and can be used in any type of Internet property, server, or network (whether it relies on <a href="http://blog.cloudflare.com/attacks-on-voip-providers/">VoIP</a> servers, UDP-based gaming servers, or HTTP servers). That said, our <a href="https://www.cloudflare.com/plans/free/">Free plan</a> can now get real-time alerts about HTTP DDoS attacks that were automatically detected and mitigated by us.</p><p>One of the benefits of Cloudflare is that all of our services and features can work together to protect your website and also improve its performance. Here’s our specialist, <a href="http://blog.cloudflare.com/author/omer/">Omer Yoachimik</a>, top 3 tips to leverage a Cloudflare free account (and put your settings more efficient to deal with DDoS attacks):</p><!--kg-card-begin: markdown--><ol>
<li>
<p>Put Cloudflare in front of your website:</p>
<ul>
<li><a href="https://developers.cloudflare.com/dns/zone-setups/full-setup/setup/">Onboard your website to Cloudflare</a> and ensure all of your HTTP traffic routes through Cloudflare. Lock down your origin server, so it only accepts traffic from <a href="https://developers.cloudflare.com/fundamentals/get-started/setup/allow-cloudflare-ip-addresses/">Cloudflare IPs</a>.</li>
</ul>
</li>
<li>
<p>Leverage Cloudflare’s free security features</p>
<ul>
<li><strong>DDoS Protection</strong>: it’s enabled by default, and if needed you can also <a href="https://developers.cloudflare.com/ddos-protection/managed-rulesets/adjust-rules/false-negative/#incomplete-mitigations">override the action to Block</a> for rules that have a different default value.</li>
<li><strong>Security Level</strong>: this feature will automatically issue challenges to requests that originate from IP addresses with low IP reputation. Ensure it's <a href="https://support.cloudflare.com/hc/en-us/articles/200170056-Understanding-the-Cloudflare-Security-Level">set to Medium</a> at least.</li>
<li><strong>Block bad bots</strong> - Cloudflare’s free tier of <a href="https://developers.cloudflare.com/bots/plans/free/">bot protection</a> can help ward off simple bots (from cloud ASNs) and headless browsers by issuing a computationally expensive challenge.</li>
<li><strong>Firewall rules</strong>: you can create up to five free <a href="https://developers.cloudflare.com/firewall/">custom firewall rules</a> to block or challenge traffic that you never want to receive.</li>
<li><strong>Managed Ruleset</strong>: in addition to your custom rule, enable Cloudflare’s <a href="https://developers.cloudflare.com/waf/managed-rulesets/">Free Managed Ruleset</a> to protect against high and wide impacting vulnerabilities</li>
</ul>
</li>
<li>
<p>Move your content to the cloud</p>
<ul>
<li><a href="https://developers.cloudflare.com/cache/">Cache</a> as much of your content as possible on the Cloudflare network. The fewer requests that hit your origin, the better — including unwanted traffic.</li>
</ul>
</li>
</ol>
<!--kg-card-end: markdown--><h2 id="2-application-level-attacks-waf">2. Application level attacks &amp; WAF</h2><p><strong>Application security: Cloudflare’s view (<a href="http://blog.cloudflare.com/application-security/">✍️</a>)</strong><br>Did you know that around 8% of all Cloudflare HTTP traffic is mitigated? That is something we explain in this application's general trends March 2022 <a href="http://blog.cloudflare.com/application-security/">blog post</a>. That means that overall, ~2.5 million requests per second are mitigated by our global network and never reach our caches or the origin servers, ensuring our customers’ bandwidth and compute power is only used for clean traffic.</p><p>You can also have a sense here of what the top mitigated traffic sources are — Layer 7 DDoS and Custom WAF (Web Application Firewall) rules are at the top — and what are the most common attacks. Other highlights include that at that time 38% of HTTP traffic we see is automated (right the number is actually lower, 31% — current trends can be seen on <a href="https://radar.cloudflare.com/">Radar</a>), and the already mentioned (about Galileo) SQLi is the most common attack vector on API endpoints.</p><p><strong>WAF for everyone: protecting the web from high severity vulnerabilities (</strong><a href="http://blog.cloudflare.com/waf-for-everyone/"><strong>✍️</strong></a><strong>)</strong><br>This <a href="http://blog.cloudflare.com/waf-for-everyone/">blog post</a> shares a relevant announcement that goes hand in hand with Cloudflare mission of "help build a better Internet" and that also includes giving some level of protection even without costs (something that also help us be better in preventing and mitigating attacks). So, since March we are providing a Cloudflare WAF Managed Ruleset that is running by default on all FREE zones, free of charge. <br><br>On this topic, there has also been a growing client side security number of threats that concerns CIOs and security professionals that we mention when we gave, in December, all paid plans access to <a href="http://blog.cloudflare.com/page-shield-generally-available/">Page Shield features</a> (last <a href="http://blog.cloudflare.com/making-page-shield-malicious-code-alerts-more-actionable/">month</a> we made Page Shield malicious code alerts more actionable. Another example is how we detect <a href="http://blog.cloudflare.com/detecting-magecart-style-attacks-for-pageshield/">Magecart-Style attacks</a> that have impacted large organizations like <a href="https://www.bbc.co.uk/news/technology-54568784">British Airways</a> and <a href="https://ico.org.uk/about-the-ico/news-and-events/news-and-blogs/2020/11/ico-fines-ticketmaster-uk-limited-125million-for-failing-to-protect-customers-payment-details/">Ticketmaster</a>, resulting in substantial GDPR fines in both cases.</p><h2 id="3-phishing-area-1-">3. Phishing (Area 1) </h2><p><strong>Why we are acquiring Area 1 (</strong><a href="http://blog.cloudflare.com/why-we-are-acquiring-area-1/"><strong>✍️</strong></a><strong>)</strong><br>Phishing remains the primary way to breach organizations. According to <a href="https://www.cisa.gov/stopransomware/general-information">CISA</a>, 90% of cyber attacks begin with it. And, in a recent report, the <a href="https://www.ic3.gov/Media/Y2022/PSA220504">FBI</a> referred to Business Email Compromise as the $43 Billion problem facing organizations.</p><p>It was in late February that it was announced that Cloudflare had agreed to acquire Area 1 Security to help organizations combat advanced email attacks and phishing campaigns. Our <a href="http://blog.cloudflare.com/why-we-are-acquiring-area-1/">blog post</a><strong> </strong>explains that “Area 1’s team has built exceptional cloud-native technology to protect businesses from email-based security threats”. So, all that technology and expertise has been integrated since then with our global network to give customers the most complete Zero Trust security platform available.<br><br><strong>The mechanics of a sophisticated phishing scam and how we stopped it (</strong><a href="http://blog.cloudflare.com/2022-07-sms-phishing-attacks/"><strong>✍️</strong></a><strong>)</strong><br>What’s in a message? Possibly a sophisticated attack targeting employees and systems. On August 8, 2022, Twilio shared that they’d been compromised by a targeted SMS phishing attack. We saw an attack with very similar characteristics also targeting Cloudflare’s employees. <a href="http://blog.cloudflare.com/2022-07-sms-phishing-attacks/">Here</a>, we do a rundown on how we were able to thwart the attack that could have breached most organizations, by using our Cloudflare One products, and physical security keys. And how others can do the same. No Cloudflare systems were compromised.</p><p>Our <a href="http://blog.cloudflare.com/introducing-cloudforce-one-threat-operations-and-threat-research/">Cloudforce One</a> threat intelligence team dissected the attack and assisted in tracking down the attacker.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image3-6.png" class="kg-image" alt="2022 attacks! An August reading list to go “Shields Up”"></figure><p><strong>Introducing browser isolation for email links to stop modern phishing threats (</strong><a href="http://blog.cloudflare.com/email-link-isolation/"><strong>✍️</strong></a><strong>)</strong><br>Why do humans <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7005690/">still click</a> on malicious links? It seems that it’s easier to do it than most people think (“human error is human”). <a href="http://blog.cloudflare.com/email-link-isolation/">Here</a> we explain how an organization nowadays can't truly have a Zero Trust security posture without securing email; an application that end users implicitly trust and threat actors take advantage of that inherent trust.</p><p>As part of our journey to integrate Area 1 into our broader Zero Trust suite, Cloudflare Gateway customers can enable Remote Browser Isolation for email links. With that, we now give unmatched level of protection from modern multi-channel email-based attacks. While we’re at it, you can also learn <a href="http://blog.cloudflare.com/replace-your-email-gateway-with-area-1/">how to replace your email gateway with Cloudflare Area 1</a>.</p><p>About <a href="http://blog.cloudflare.com/account-compromise-security-overview/"><strong>account takeovers</strong></a>, we explained back in March 2021 how we prevent account takeovers on our own applications (on the phishing side we were already using, as a customer, at the time, Area 1).</p><p>Also from last year, <a href="http://blog.cloudflare.com/research-directions-in-password-security/">here’s</a> our research in <strong>password security </strong>(and the problem of password reuse) — it gets technical. There’s a new password related protocol called OPAQUE (<a href="https://opaque-full.research.cloudflare.com/">we added a new demo about it on January 2022</a>) that could help better store secrets that our research team is excited about.</p><h2 id="4-malware-ransomware-other-risks">4. Malware/Ransomware &amp; other risks</h2><p><strong>How Cloudflare Security does Zero Trust (</strong><a href="http://blog.cloudflare.com/how-cloudflare-security-does-zero-trust/"><strong>✍️</strong></a><strong>)</strong><br>Security is more than ever part of an ecosystem that the more robust, the more efficient in avoiding or mitigating attacks. In this <a href="http://blog.cloudflare.com/how-cloudflare-security-does-zero-trust/">blog post</a> written for our <a href="https://www.cloudflare.com/cloudflare-one-week/">Cloudflare One week</a>, we explain how that ecosystem, in this case inside our Zero Trust services, can give protection from malware, ransomware, phishing, command &amp; control, shadow IT, and other Internet risks over all ports and protocols.</p><p>Since 2020, we launched <a href="http://blog.cloudflare.com/announcing-antivirus-in-cloudflare-gateway/">Cloudflare Gateway</a> focused on malware detection and prevention directly from the Cloudflare edge. Recently, we also include our new <a href="https://www.cloudflare.com/products/zero-trust/casb/">CASB</a> product (to secure workplace tools, personalize access, secure sensitive data).</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image1-14.png" class="kg-image" alt="2022 attacks! An August reading list to go “Shields Up”"></figure><p><strong>Anatomy of a Targeted Ransomware Attack (</strong><a href="http://blog.cloudflare.com/targeted-ransomware-attack/"><strong>✍️</strong></a><strong>)</strong><br>What a ransomware attack looks like for the victim:</p><blockquote><em>“Imagine your most critical systems suddenly stop operating. And then someone demands a ransom to get your systems working again. Or someone launches a DDoS against you and demands a ransom to make it stop. That’s the world of ransomware and ransom DDoS.”</em></blockquote><p>Ransomware attacks continue <a href="https://www.kroll.com/en/insights/publications/cyber/ransomware-attack-trends-2020">to be on the rise</a> and there’s no sign of them slowing down in the near future. That was true more than a year ago, when this <a href="http://blog.cloudflare.com/targeted-ransomware-attack/">blog post</a> was written and is still <a href="https://www.fitchratings.com/research/corporate-finance/ransomware-growing-cyber-risk-for-us-corporates-financials-govt-27-04-2022">ongoing</a>, up 105% YoY according to a Senate Committee March 2022 report. And the nature of ransomware attacks is changing. Here, we highlight how <a href="https://www.cloudflare.com/learning/ddos/ransom-ddos-attack/">Ransom DDoS (RDDoS)</a> attacks work, how Cloudflare onboarded and <a href="http://blog.cloudflare.com/ransom-ddos-attacks-target-a-fortune-global-500-company/">protected</a> a Fortune 500 customer from a targeted one, and how that <a href="http://blog.cloudflare.com/announcing-antivirus-in-cloudflare-gateway/">Gateway with antivirus</a> we mentioned before helps with just that.</p><p>We also show that with ransomware as a service (<a href="https://www.crowdstrike.com/cybersecurity-101/ransomware/ransomware-as-a-service-raas/">RaaS</a>) models, it’s even easier for inexperienced threat actors to get their hands on them today (“RaaS is essentially a franchise that allows criminals to rent ransomware from malware authors”). We also include some general recommendations to help you and your organization stay secure. Don’t want to click the link? Here they are:</p><ul><li>Use 2FA everywhere, especially on your remote access entry points. This is where Cloudflare Access really helps.</li><li>Maintain multiple redundant backups of critical systems and data, both onsite and offsite</li><li>Monitor and block malicious domains using Cloudflare Gateway + AV</li><li>Sandbox web browsing activity using Cloudflare RBI to isolate threats at the browser</li></ul><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image2-7.png" class="kg-image" alt="2022 attacks! An August reading list to go “Shields Up”"></figure><p><strong>Investigating threats using the Cloudflare Security Center (</strong><a href="http://blog.cloudflare.com/security-center-investigate/"><strong>✍️</strong></a><strong>)</strong><br><a href="http://blog.cloudflare.com/security-center-investigate/">Here</a>, first we announce our new threat investigations portal, <em>Investigate</em>, right in the Cloudflare Security Center, that allows all customers to query directly our intelligence to streamline security workflows and tighten feedback loops.</p><p>That’s only possible because we have a global and in-depth view, given that we protect millions of Internet properties from attacks (the free plans help us to have that insight). And the data we glean from these attacks trains our machine learning models and improves the efficacy of our network and application security products.</p><p><strong>Steps we've taken around Cloudflare's services in Ukraine, Belarus, and Russia (</strong><a href="http://blog.cloudflare.com/steps-taken-around-cloudflares-services-in-ukraine-belarus-and-russia/"><strong>✍️</strong></a><strong>)</strong><br>There’s an emergence of the known as <a href="https://en.wikipedia.org/wiki/Wiper_(malware)">wiper</a> malware attacks (intended to erase the computer it infects) and in this <a href="http://blog.cloudflare.com/steps-taken-around-cloudflares-services-in-ukraine-belarus-and-russia/">blog post</a>, among other things, we explain how when a wiper malware was identified in Ukraine (it took offline government agencies and a major bank), we successfully adapted our Zero Trust products to make sure our customers were protected. Those protections include many Ukrainian organizations, under our <a href="http://blog.cloudflare.com/in-ukraine-and-beyond-what-it-takes-to-keep-vulnerable-groups-online/">Project Galileo</a> that is having a busy year, and they were automatically put available to all our customers. More recently, the satellite provider Viasat was <a href="https://techcrunch.com/2022/05/10/russia-viasat-cyberattack/">affected</a>.</p><p><strong>Zaraz use Workers to make third-party tools secure and fast (</strong><a href="http://blog.cloudflare.com/zaraz-use-workers-to-make-third-party-tools-secure-and-fast/"><strong>✍️</strong></a><strong>)</strong><br>Cloudflare announced it acquired <a href="http://blog.cloudflare.com/cloudflare-acquires-zaraz-to-enable-cloud-loading-of-third-party-tools/">Zaraz</a> in December 2021 to help us enable cloud loading of third-party tools. Seems unrelated to attacks? Think again (this takes us back to the secure ecosystem I already mentioned). Among other things, <a href="http://blog.cloudflare.com/zaraz-use-workers-to-make-third-party-tools-secure-and-fast/"><strong>here</strong></a> you can learn how Zaraz can make your website more secure (and faster) by offloading third-party scripts.</p><p>That allows to avoid problems and attacks. Which? From code tampering to lose control over the data sent to third-parties. My colleague <a href="http://blog.cloudflare.com/author/yoav/">Yo'av Moshe</a> elaborates on what this solution prevents: “the third-party script can intentionally or unintentionally (due to being hacked) collect information it shouldn't collect, like credit card numbers, Personal Identifiers Information (PIIs), etc.”. You should definitely avoid those.</p><p><strong>Introducing Cloudforce One: our new threat operations and research team (</strong><a href="http://blog.cloudflare.com/introducing-cloudforce-one-threat-operations-and-threat-research/"><strong>✍️</strong></a><strong>)</strong><br><a href="http://blog.cloudflare.com/introducing-cloudforce-one-threat-operations-and-threat-research/">Meet</a> our new threat operations and research team: <strong>Cloudforce One</strong>. While this team will publish research, that’s not its reason for being. Its primary objective: track and disrupt threat actors. It’s all about being protected against a great flow of threats with minimal to no involvement.</p><h2 id="wrap-up">Wrap up</h2><p>The expression “if it ain't broke, don't fix it” doesn’t seem to apply to the fast pacing Internet industry, where attacks are also in the fast track. If you or your company and services aren’t properly protected, attackers (human or bots) will probably find you sooner than later (maybe they already did).</p><p>To end on a popular quote used in books, movies and in life: “You keep knocking on the devil's door long enough and sooner or later someone's going to answer you”. Although we have been onboarding many organizations while attacks are happening, that’s not the less hurtful solution — preventing and mitigating effectively and forget the protection is even there.</p><p>If you want to try some security features mentioned, the <a href="https://www.cloudflare.com/securitycenter/">Cloudflare Security Center</a> is a good place to start (free plans included). The same with our <a href="https://www.cloudflare.com/plans/zero-trust-services/">Zero Trust ecosystem</a> (or <a href="https://www.cloudflare.com/cloudflare-one/">Cloudflare One</a> as our SASE, Secure Access Service Edge) that is available as self-serve, and also includes a free plan (this vendor-agnostic <a href="https://zerotrustroadmap.org/">roadmap</a> shows the general advantages of the Zero Trust architecture).</p><p>If trends are more your thing, <a href="https://radar.cloudflare.com/">Cloudflare Radar</a> has a near real-time dedicated area about attacks, and you can browse and interact with our <a href="https://radar.cloudflare.com/notebooks/ddos-2022-q2">DDoS attack trends for 2022 Q2</a> report.</p>]]></content:encoded></item><item><title><![CDATA[The mechanics of a sophisticated phishing scam and how we stopped it]]></title><description><![CDATA[Yesterday, August 8, 2022, Twilio shared that they’d been compromised by a targeted phishing attack. Around the same time as Twilio was attacked, we saw an attack with very similar characteristics also targeting Cloudflare’s employees]]></description><link>https://blog.cloudflare.com/2022-07-sms-phishing-attacks/</link><guid isPermaLink="false">62f27978bc3891000abd9426</guid><category><![CDATA[Security]]></category><category><![CDATA[Post Mortem]]></category><category><![CDATA[Phishing]]></category><category><![CDATA[Cloudflare Gateway]]></category><category><![CDATA[Cloudflare Access]]></category><dc:creator><![CDATA[Matthew Prince]]></dc:creator><pubDate>Tue, 09 Aug 2022 15:56:30 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/unnamed.png" medium="image"/><content:encoded><![CDATA[<!--kg-card-begin: markdown--><img src="http://blog.cloudflare.com/content/images/2022/08/unnamed.png" alt="The mechanics of a sophisticated phishing scam and how we stopped it"><p><em><small>This post is also available in <a href="http://blog.cloudflare.com/zh-cn/2022-07-sms-phishing-attacks-zh-cn/">简体中文</a>, <a href="http://blog.cloudflare.com/ja-jp/2022-07-sms-phishing-attacks-ja-jp/">日本語</a> and <a href="http://blog.cloudflare.com/es-es/2022-07-sms-phishing-attacks-es-es/">Español</a>.</small></em></p>
<!--kg-card-end: markdown--><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/unnamed-1.png" class="kg-image" alt="The mechanics of a sophisticated phishing scam and how we stopped it"></figure><p>Yesterday, August 8, 2022, Twilio shared that they’d been <a href="https://www.twilio.com/blog/august-2022-social-engineering-attack">compromised by a targeted phishing attack</a>. Around the same time as Twilio was attacked, we saw an attack with very similar characteristics also targeting Cloudflare’s employees. While individual employees did fall for the phishing messages, we were able to thwart the attack through our own use of <a href="https://www.cloudflare.com/cloudflare-one/">Cloudflare One products</a>, and physical security keys issued to every employee that are required to access all our applications.</p><p>We have confirmed that no Cloudflare systems were compromised. Our <a href="http://blog.cloudflare.com/introducing-cloudforce-one-threat-operations-and-threat-research/">Cloudforce One threat intelligence team</a> was able to perform additional analysis to further dissect the mechanism of the attack and gather critical evidence to assist in tracking down the attacker.</p><p>This was a sophisticated attack targeting employees and systems in such a way that we believe most organizations would be likely to be breached. Given that the attacker is targeting multiple organizations, we wanted to share here a rundown of exactly what we saw in order to help other companies recognize and mitigate this attack.</p><h2 id="targeted-text-messages">Targeted Text Messages</h2><p>On July 20, 2022, the Cloudflare Security team received reports of employees receiving legitimate-looking text messages pointing to what appeared to be a Cloudflare Okta login page. The messages began at 2022-07-20 22:50 UTC. Over the course of less than 1 minute, at least 76 employees received text messages on their personal and work phones. Some messages were also sent to the employees family members. We have not yet been able to determine how the attacker assembled the list of employees phone numbers but have reviewed access logs to our employee directory services and have found no sign of compromise.</p><p>Cloudflare runs a 24x7 Security Incident Response Team (SIRT). Every Cloudflare employee is trained to report anything that is suspicious to the SIRT. More than 90 percent of the reports to SIRT turn out to not be threats. Employees are encouraged to report anything and never discouraged from over-reporting. In this case, however, the reports to SIRT were a real threat.</p><p>The text messages received by employees looked like this:</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image3-5.png" class="kg-image" alt="The mechanics of a sophisticated phishing scam and how we stopped it"></figure><p>They came from four phone numbers associated with T-Mobile-issued SIM cards: (754) 268-9387, (205) 946-7573, (754) 364-6683 and (561) 524-5989. They pointed to an official-looking domain: cloudflare-okta.com. That domain had been registered via Porkbun, a domain registrar, at 2022-07-20 22:13:04 UTC — less than 40 minutes before the phishing campaign began.</p><p>Cloudflare built our <a href="https://www.cloudflare.com/products/registrar/custom-domain-protection/">secure registrar product</a> in part to be able to monitor when domains using the Cloudflare brand were registered and get them shut down. However, because this domain was registered so recently, it had not yet been published as a new .com registration, so our systems did not detect its registration and our team had not yet moved to terminate it.</p><p>If you clicked on the link it took you to a phishing page. The phishing page was hosted on DigitalOcean and looked like this:</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image1-13.png" class="kg-image" alt="The mechanics of a sophisticated phishing scam and how we stopped it"></figure><p>Cloudflare uses Okta as our identity provider. The phishing page was designed to look identical to a legitimate Okta login page. The phishing page prompted anyone who visited it for their username and password.</p><h2 id="real-time-phishing">Real-Time Phishing</h2><p>We were able to analyze the payload of the <a href="https://www.cloudflare.com/learning/email-security/what-is-email-security/">phishing attack</a> based on what our employees received as well as its content being posted to services like VirusTotal by other companies that had been attacked. When the phishing page was completed by a victim, the credentials were immediately relayed to the attacker via the messaging service Telegram. This real-time relay was important because the phishing page would also prompt for a Time-based One Time Password (TOTP) code.</p><p>Presumably, the attacker would receive the credentials in real-time, enter them in a victim company’s actual login page, and, for many organizations that would generate a code sent to the employee via SMS or displayed on a password generator. The employee would then enter the TOTP code on the phishing site, and it too would be relayed to the attacker. The attacker could then, before the TOTP code expired, use it to access the company’s actual login page — defeating most two-factor authentication implementations.</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image2-6.png" class="kg-image" alt="The mechanics of a sophisticated phishing scam and how we stopped it"></figure><h2 id="protected-even-if-not-perfect">Protected Even If Not Perfect</h2><p>We confirmed that three Cloudflare employees fell for the phishing message and entered their credentials. However, Cloudflare does not use TOTP codes. Instead, every employee at the company is issued a FIDO2-compliant security key from a vendor like YubiKey. Since the hard keys are tied to users and implement <a href="https://www.yubico.com/blog/creating-unphishable-security-key/">origin binding</a>, even a sophisticated, real-time phishing operation like this cannot gather the information necessary to log in to any of our systems. While the attacker attempted to log in to our systems with the compromised username and password credentials, they could not get past the hard key requirement.</p><p>But this phishing page was not simply after credentials and TOTP codes. If someone made it past those steps, the phishing page then initiated the download of a phishing payload which included AnyDesk’s remote access software. That software, if installed, would allow an attacker to control the victim’s machine remotely. We confirmed that none of our team members got to this step. If they had, however, our endpoint security would have stopped the installation of the remote access software.</p><h2 id="how-did-we-respond">How Did We Respond?</h2><p>The main response actions we took for this incident were:</p><h3 id="1-block-the-phishing-domain-using-cloudflare-gateway">1. Block the phishing domain using Cloudflare Gateway</h3><p>Cloudflare Gateway is a <a href="https://www.cloudflare.com/learning/access-management/what-is-a-secure-web-gateway/">Secure Web Gateway</a> solution providing threat and data protection with DNS / HTTP filtering and natively-integrated <a href="https://www.cloudflare.com/learning/security/glossary/what-is-zero-trust/">Zero Trust</a>. We use this  solution internally to proactively identify malicious domains and block them. Our team added the malicious domain to Cloudflare Gateway to block all employees from accessing it.</p><p>Gateway’s automatic detection of malicious domains also identified the domain and blocked it, but the fact that it was registered and messages were sent within such a short interval of time meant that the system hadn’t automatically taken action before some employees had clicked on the links. Given this incident we are working to speed up how quickly malicious domains are identified and blocked. We’re also implementing controls on access to newly registered domains which we offer to customers but had not implemented ourselves.</p><h3 id="2-identify-all-impacted-cloudflare-employees-and-reset-compromised-credentials">2. Identify all impacted Cloudflare employees and reset compromised credentials</h3><p>We were able to compare recipients of the phishing texts to login activity and identify threat-actor attempts to authenticate to our employee accounts. We identified login attempts blocked due to the hard key (U2F) requirements indicating that the correct password was used, but the second factor could not be verified. For the three of our employees' credentials were leaked, we reset their credentials and any active sessions and initiated scans of their devices.</p><h3 id="3-identify-and-take-down-threat-actor-infrastructure">3. Identify and take down threat-actor infrastructure</h3><p>The threat actor's phishing domain was newly registered via Porkbun, and hosted on DigitalOcean. The phishing domain used to target Cloudflare was set up less than an hour before the initial phishing wave. The site had a Nuxt.js frontend, and a Django backend. We worked with DigitalOcean to shut down the attacker’s server. We also worked with Porkbun to seize control of the malicious domain.</p><p>From the failed sign-in attempts we were able to determine that the threat actor was leveraging Mullvad VPN software and distinctively using the Google Chrome browser on a Windows 10 machine. The VPN IP addresses used by the attacker were 198.54.132.88, and 198.54.135.222. Those IPs are assigned to Tzulo, a US-based dedicated server provider whose website claims they have servers located in Los Angeles and Chicago. It appears, actually, that the first was actually running on a server in the Toronto area and the latter on a server in the Washington, DC area. We blocked these IPs from accessing any of our services.</p><h3 id="4-update-detections-to-identify-any-subsequent-attack-attempts">4. Update detections to identify any subsequent attack attempts</h3><p>With what we were able to uncover about this attack, we incorporated additional signals to our already existing detections to specifically identify this threat-actor. At the time of writing we have not observed any additional waves targeting our employees. However, intelligence from the server indicated the attacker was targeting other organizations, including Twilio. We reached out to these other organizations and shared intelligence on the attack.</p><h3 id="5-audit-service-access-logs-for-any-additional-indications-of-attack">5. Audit service access logs for any additional indications of attack</h3><p>Following the attack, we screened all our system logs for any additional fingerprints from this particular attacker. Given Cloudflare Access serves as the central control point for all Cloudflare applications, we can search the logs for any indication the attacker may have breached any systems. Given employees’ phones were targeted, we also carefully reviewed the logs of our employee directory providers. We did not find any evidence of compromise.</p><h2 id="lessons-learned-and-additional-steps-we-re-taking">Lessons Learned and Additional Steps We’re Taking</h2><p>We learn from every attack. Even though the attacker was not successful, we are making additional adjustments from what we’ve learned. We’re adjusting the settings for Cloudflare Gateway to restrict or sandbox access to sites running on domains that were registered within the last 24 hours. We will also run any non-whitelisted sites containing terms such as “cloudflare” “okta” “sso” and “2fa” through our browser isolation technology. We are also increasingly using <a href="https://www.cloudflare.com/products/zero-trust/email-security/">Cloudflare Area 1’s phish-identification technology</a> to scan the web and look for any pages that are designed to target Cloudflare. Finally, we’re tightening up our Access implementation to prevent any logins from unknown VPNs, residential proxies, and infrastructure providers. All of these are standard features of the same products we offer to customers.</p><p>The attack also reinforced the importance of three things we’re doing well. First, requiring hard keys for access to all applications. <a href="https://krebsonsecurity.com/2018/07/google-security-keys-neutralized-employee-phishing/">Like Google</a>, we have not seen any successful phishing attacks since rolling hard keys out. Tools like Cloudflare Access made it easy to support hard keys even across legacy applications. If you’re an organization interested in how we rolled out hard keys, reach out to <a href="mailto:cloudforceone-irhelp@cloudflare.com">cloudforceone-irhelp@cloudflare.com</a> and our security team would be happy to share the best practices we learned through this process.</p><p>Second, using Cloudflare’s own technology to protect our employees and systems. Cloudflare One’s solutions like Access and Gateway were critical to staying ahead of this attack. We configured our Access implementation to require hard keys for every application. It also creates a central logging location for all application authentications. And, if ever necessary, a place from which we can kill the sessions of a potentially compromised employee. Gateway allows us the ability to shut down malicious sites like this one quickly and understand what employees may have fallen for the attack. These are all functionalities that we make available to Cloudflare customers as part of our Cloudflare One suite and this attack demonstrates how effective they can be.</p><p>Third, having a paranoid but blame-free culture is critical for security. The three employees who fell for the phishing scam were not reprimanded. We’re all human and we make mistakes. It’s critically important that when we do, we report them and don’t cover them up. This incident provided another example of why security is part of every team member at Cloudflare’s job.</p><h2 id="detailed-timeline-of-events">Detailed Timeline of Events</h2><!--kg-card-begin: markdown--><style type="text/css">
.tg  {border-collapse:collapse;border-spacing:0;}
.tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;
  overflow:hidden;padding:10px 5px;word-break:normal;}
.tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;
  font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;}
.tg .tg-0lax{text-align:left;vertical-align:top}
</style>
<table class="tg">
<thead>
  <tr>
    <th class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">2022-07-20 22:49 UTC</span></th>
    <th class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Attacker sends out 100+ SMS messages to Cloudflare employees and their families.</span></th>
  </tr>
</thead>
<tbody>
  <tr>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">2022-07-20 22:50 UTC</span></td>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Employees begin reporting SMS messages to Cloudflare Security team.</span></td>
  </tr>
  <tr>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">2022-07-20 22:52 UTC</span></td>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Verify that the attacker's domain is blocked in Cloudflare Gateway for corporate devices.</span></td>
  </tr>
  <tr>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">2022-07-20 22:58 UTC</span></td>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Warning communication sent to all employees across chat and email.</span></td>
  </tr>
  <tr>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">2022-07-20 22:50 UTC to</span><br><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">2022-07-20 23:26 UTC</span></td>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Monitor telemetry in the Okta System log &amp; Cloudflare Gateway HTTP logs to locate credential compromise. Clear login sessions and suspend accounts on discovery.</span></td>
  </tr>
  <tr>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">2022-07-20 23:26 UTC</span></td>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Phishing site is taken down by the hosting provider.</span></td>
  </tr>
  <tr>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">2022-07-20 23:37 UTC</span></td>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Reset leaked employee credentials. </span></td>
  </tr>
  <tr>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">2022-07-21 00:15 UTC</span></td>
    <td class="tg-0lax"><span style="font-weight:400;font-style:normal;text-decoration:none;color:#000;background-color:transparent">Deep dive into attacker infrastructure and capabilities.</span></td>
  </tr>
</tbody>
</table><!--kg-card-end: markdown--><h2 id="indicators-of-compromise">Indicators of compromise</h2><!--kg-card-begin: html--><style type="text/css">
.tg  {border-collapse:collapse;border-spacing:0;}
.tg td{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;
  overflow:hidden;padding:10px 5px;word-break:normal;}
.tg th{border-color:black;border-style:solid;border-width:1px;font-family:Arial, sans-serif;font-size:14px;
  font-weight:normal;overflow:hidden;padding:10px 5px;word-break:normal;}
.tg .tg-nr0u{border-color:inherit;font-family:inherit;font-size:100%;text-align:left;vertical-align:top}
.tg .tg-0pky{border-color:inherit;text-align:left;vertical-align:top}
</style>
<table class="tg">
<thead>
  <tr>
    <th class="tg-nr0u">Value</th>
    <th class="tg-0pky">Type</th>
    <th class="tg-0pky">Context and MITRE Mapping</th>
  </tr>
</thead>
<tbody>
  <tr>
    <td class="tg-0pky">cloudflare-okta[.]com hosted on 147[.]182[.]132[.]52</td>
    <td class="tg-0pky">Phishing URL</td>
    <td class="tg-0pky"><a href="https://attack.mitre.org/techniques/T1566/002/">T1566.002</a>: Phishing: Spear Phishing Link sent to users.</td>
  </tr>
  <tr>
    <td class="tg-0pky">64547b7a4a9de8af79ff0eefadde2aed10c17f9d8f9a2465c0110c848d85317a</td>
    <td class="tg-0pky">SHA-256</td>
    <td class="tg-0pky"><a href="https://attack.mitre.org/techniques/T1219/">T1219</a>: Remote Access Software being distributed by the threat actor</td>
  </tr>
</tbody>
</table><!--kg-card-end: html--><h2 id="what-you-can-do">What You Can Do</h2><p>If you are seeing similar attacks in your environment, please don’t hesitate to reach out to <a href="mailto:cloudforceone-irhelp@cloudflare.com">cloudforceone-irhelp@cloudflare.com</a>, and we’re happy to share best practices on how to keep your business secure. Finally, do you want to work on detecting and mitigating the next attacks with us? We’re hiring on our Detection and Response team, <a href="https://boards.greenhouse.io/cloudflare/jobs/4364485?gh_jid=4364485">come join us</a>!</p>]]></content:encoded></item><item><title><![CDATA[Introducing new Cloudflare for SaaS documentation]]></title><description><![CDATA[Cloudflare for SaaS offers a suite of Cloudflare products and add-ons to improve the security, performance, and reliability of SaaS providers. Now, the Cloudflare for SaaS documentation outlines how to optimize it in order to meet your goals]]></description><link>https://blog.cloudflare.com/introducing-new-cloudflare-for-saas-documentation/</link><guid isPermaLink="false">62f0cad5bc3891000abd93d1</guid><category><![CDATA[Technical Writing]]></category><category><![CDATA[Developer Documentation]]></category><category><![CDATA[Cloudflare for SaaS]]></category><category><![CDATA[SSL]]></category><category><![CDATA[SaaS]]></category><dc:creator><![CDATA[Mia Malden]]></dc:creator><pubDate>Tue, 09 Aug 2022 13:00:00 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/image3-3.png" medium="image"/><content:encoded><![CDATA[<figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image3-4.png" class="kg-image" alt="Introducing new Cloudflare for SaaS documentation"></figure><img src="http://blog.cloudflare.com/content/images/2022/08/image3-3.png" alt="Introducing new Cloudflare for SaaS documentation"><p>As a SaaS provider, you’re juggling many challenges while building your application, whether it’s custom domain support, protection from attacks, or maintaining an origin server. In 2021, we were proud to announce <a href="http://blog.cloudflare.com/cloudflare-for-saas/">Cloudflare for SaaS for Everyone</a>, which allows anyone to use Cloudflare to cover those challenges, so they can focus on other aspects of their business. This product has a variety of potential implementations; now, we are excited to announce a new section in our <a href="https://developers.cloudflare.com/">Developer Docs</a> specifically devoted to <a href="https://developers.cloudflare.com/cloudflare-for-saas/">Cloudflare for SaaS documentation</a> to allow you take full advantage of its product suite.</p><h3 id="cloudflare-for-saas-solution">Cloudflare for SaaS solution</h3><p>You may remember, from our <a href="http://blog.cloudflare.com/cloudflare-for-saas-for-all-now-generally-available/">October 2021 blog post</a>, all the ways that Cloudflare provides solutions for SaaS providers:</p><ul><li>Set up an origin server</li><li>Encrypt your customers’ traffic</li><li>Keep your customers online</li><li>Boost the performance of global customers</li><li>Support custom domains</li><li>Protect against attacks and bots</li><li>Scale for growth</li><li>Provide insights and analytics</li></ul><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image2-5.png" class="kg-image" alt="Introducing new Cloudflare for SaaS documentation"></figure><p>However, we received feedback from customers indicating confusion around actually <em>using</em> the capabilities of Cloudflare for SaaS because there are so many features! With the existing documentation, it wasn’t 100% clear how to enhance security and performance, or how to support custom domains. Now, we want to show customers how to use Cloudflare for SaaS to its full potential by including more product integrations in the docs, as opposed to only focusing on the SSL/TLS piece.</p><h3 id="bridging-the-gap">Bridging the gap</h3><p>Cloudflare for SaaS can be overwhelming with so many possible add-ons and configurations. That’s why the new docs are organized into six main categories, housing a number of new, detailed guides (for example, WAF for SaaS and Regional Services for SaaS):</p><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/image1-12.png" class="kg-image" alt="Introducing new Cloudflare for SaaS documentation"></figure><p>Once you get your SaaS application up and running with the <a href="https://developers.cloudflare.com/cloudflare-for-saas/getting-started/">Get Started</a> page, you can find which configurations are best suited to your needs based on your priorities as a provider. Even if you aren’t sure what your goals are, this setup outlines the possibilities much more clearly through a number of new documents and product guides such as:</p><ul><li><a href="https://developers.cloudflare.com/cloudflare-for-saas/start/advanced-settings/regional-services-for-saas/">Regional Services for SaaS</a></li><li><a href="https://developers.cloudflare.com/analytics/graphql-api/tutorials/end-customer-analytics/">Querying HTTP events by hostname with GraphQL</a></li><li><a href="https://developers.cloudflare.com/cloudflare-for-saas/domain-support/migrating-custom-hostnames/">Migrating custom hostnames</a></li></ul><p>Instead of pondering over vague subsection titles, you can peruse with purpose in mind. The advantages and possibilities of Cloudflare for SaaS are highlighted instead of hidden.</p><h3 id="possible-configurations">Possible configurations</h3><p>This setup facilitates configurations much more easily to meet your goals as a SaaS provider.</p><p>For example, consider performance. Previously, there was no documentation surrounding reduced latency for SaaS providers. Now, the Performance section explains the automatic benefits to your performance by onboarding with Cloudflare for SaaS. Additionally, it offers three options of how to reduce latency even further through brand-new docs:</p><ul><li><a href="https://developers.cloudflare.com/cloudflare-for-saas/performance/early-hints-for-saas/">Early Hints for SaaS</a></li><li><a href="https://developers.cloudflare.com/cloudflare-for-saas/performance/cache-for-saas/">Cache for SaaS</a></li><li><a href="https://developers.cloudflare.com/cloudflare-for-saas/performance/argo-for-saas/">Argo Smart Routing for SaaS</a></li></ul><p>Similarly, the new organization offers <a href="https://developers.cloudflare.com/cloudflare-for-saas/security/waf-for-saas/">WAF for SaaS</a> as a previously hidden security solution, extending providers the ability to enable automatic protection from vulnerabilities and the flexibility to create custom rules. This is conveniently accompanied by a <a href="https://developers.cloudflare.com/cloudflare-for-saas/security/waf-for-saas/managed-rulesets/">step-by-step tutorial using Cloudflare Managed Rulesets</a>.</p><h3 id="what-s-next">What’s next</h3><p>While this transition represents an improvement in the Cloudflare for SaaS docs, we’re going to expand its accessibility even more. Some tutorials, such as our <a href="https://developers.cloudflare.com/cloudflare-for-saas/security/waf-for-saas/managed-rulesets/">Managed Ruleset Tutorial</a>, are already live within the tile. However, more step-by-step guides for Cloudflare for SaaS products and add-ons will further enable our customers to take full advantage of the available product suite. In particular, keep an eye out for expanding documentation around using Workers for Platforms.</p><h3 id="check-it-out">Check it out</h3><p>Visit the new <a href="http://www.developers.cloudflare.com/cloudflare-for-saas">Cloudflare for SaaS tile</a> to see the updates. If you are a SaaS provider interested in extending Cloudflare benefits to your customers through Cloudflare for SaaS, visit our <a href="https://www.cloudflare.com/saas/">Cloudflare for SaaS overview</a> and our <a href="https://developers.cloudflare.com/cloudflare-for-saas/plans/">Plans page</a>.</p>]]></content:encoded></item><item><title><![CDATA[1.1.1.1 + WARP: More features, still private]]></title><description><![CDATA[We’re announcing two major improvements to our 1.1.1.1 + WARP apps]]></description><link>https://blog.cloudflare.com/geoexit-improving-warp-user-experience-larger-network/</link><guid isPermaLink="false">62e0ffcc1768f8000a2ae4e1</guid><category><![CDATA[WARP]]></category><category><![CDATA[Privacy]]></category><category><![CDATA[1.1.1.1]]></category><category><![CDATA[Security]]></category><dc:creator><![CDATA[Mari Galicer]]></dc:creator><pubDate>Sat, 06 Aug 2022 16:15:03 GMT</pubDate><media:content url="http://blog.cloudflare.com/content/images/2022/08/WARP-Geoexit-1.png" medium="image"/><content:encoded><![CDATA[<figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/08/WARP-Geoexit.png" class="kg-image" alt="1.1.1.1 + WARP: More features, still private"></figure><img src="http://blog.cloudflare.com/content/images/2022/08/WARP-Geoexit-1.png" alt="1.1.1.1 + WARP: More features, still private"><p>It’s a Saturday night. You open your browser, looking for nearby pizza spots that are open. If the search goes as intended, your browser will show you results that are within a few miles, often based on the assumed location of your IP address. At Cloudflare, we affectionately call this type of geolocation accuracy the “pizza test”. When you use a Cloudflare product that sits between you and the Internet (for example, <a href="http://blog.cloudflare.com/1111-warp-better-vpn/">WARP</a>), it’s one of the ways we work to balance user experience and privacy. Too inaccurate and you’re getting pizza places from a neighboring country; too accurate and you’re reducing the privacy benefits of obscuring your location.</p><p>With that in mind, we’re excited to announce two major improvements to our 1.1.1.1 + WARP apps: first, an improvement to how we ensure search results and other geographically-aware Internet activity work without compromising your privacy, and second, a larger network with more locations available to WARP+ subscribers, powering even speedier connections to our global network.</p><h3 id="a-better-internet-browsing-experience-for-every-warp-user">A better Internet browsing experience for every WARP user</h3><p>When we originally built the 1.1.1.1+ WARP mobile app, we wanted to create a consumer-friendly way to connect to our network and privacy-respecting <a href="https://1.1.1.1/">DNS resolver</a>.</p><p>What we discovered over time is that the topology of the Internet dictates a different type of experience for users in different locations. Why? Sometimes, because traffic congestion or technical issues route your traffic to a less congested part of the network. Other times, Internet Service Providers may not <a href="https://www.cloudflare.com/peering-policy/">peer with Cloudflare</a> or engage in traffic engineering to optimize their networks how they see fit, which could result in user traffic connecting to a location that doesn’t quite map to their locale or language.</p><p>Regardless of the cause, the impact is that your search results become less relevant, if not outright confusing. For example, in somewhere dense with country borders, like Europe, your traffic in Berlin could get routed to Amsterdam because your mobile operator chooses to not peer in-country, giving you results in Dutch instead of German. This can also be disruptive if you’re trying to stream content subject to licensing restrictions, such as a person in the UK trying to watch BBC iPlayer or a person in Brazil trying to watch the World Cup.</p><p>So we fixed this. We just rolled out a major update to the service that powers WARP that will give you a geographically accurate browsing experience without revealing your IP address to the websites you’re visiting. Instead, websites you visit will see a Cloudflare IP address instead, making it harder for them to track you directly.</p><h3 id="how-it-works">How it works</h3><p>Traditionally, consumer VPNs deliberately route your traffic through a server in another country, making your connection slow, and often getting blocked because of their ability to flout location-based content restrictions. We took a different approach when we first launched WARP in 2018, giving you the best possible performance by routing your traffic through the Cloudflare data center closest to you. However, because not every Internet Service Provider (ISP) peers with Cloudflare, users sometimes end up exiting the Cloudflare network from a more “random” data center – one that does not accurately represent their locale.</p><p>Websites and third party services often infer geolocation from your IP address, and now, 1.1.1.1 + WARP replaces your original IP address with one that consistently and accurately represents your approximate location.</p><p>Here’s how we did it:</p><ol><li>We ran an analysis on a subset of our network traffic to find a rough approximation of how many users we have per city.</li><li>We divided that amongst our egress IPs, using an anycast architecture to be efficient with the number of additional IPs we had to allocate and advertise per metro area.</li><li>We then submitted geolocation information of those IPs to various geolocation database providers, ensuring third party services associate those Cloudflare egress IPs with an accurate approximate location.</li></ol><p>It was important to us to provide the benefits of this location accuracy without compromising user privacy, so the app doesn’t ask for specific location permissions or log your IP address.</p><figure class="kg-card kg-image-card kg-width-wide"><img src="http://blog.cloudflare.com/content/images/2022/07/image1-12.png" class="kg-image" alt="1.1.1.1 + WARP: More features, still private"></figure><h3 id="an-even-bigger-network-for-warp-users">An even bigger network for WARP+ users</h3><figure class="kg-card kg-image-card"><img src="http://blog.cloudflare.com/content/images/2022/07/image2-22.png" class="kg-image" alt="1.1.1.1 + WARP: More features, still private"></figure><p>We also recently announced that we’ve expanded our network to <a href="http://blog.cloudflare.com/new-cities-april-2022-edition/">over 275 cities</a> in over 100 countries. This gave us an opportunity to revisit where we offered WARP, and how we could expand the number of locations users can connect to WARP with (in other words: an opportunity to make things faster).</p><p>From today, all WARP+ subscribers will benefit from a larger network with 20+ new cities: with no change in subscription pricing. A closer Cloudflare data center means less latency between your device and Cloudflare, which directly improves your download speed, thanks to what’s called the <a href="https://en.wikipedia.org/wiki/Bandwidth-delay_product">Bandwidth-Delay Product</a> (put simply: lower latency, higher throughput!).</p><p>As a result, sites load faster, both for those on the <a href="https://www.cloudflare.com/network/">Cloudflare network</a> and those that aren’t. As we continue to expand our network, we'll be revising this on a regular basis to ensure that all WARP and WARP+ subscribers continue to get great performance.</p><h3 id="speed-privacy-and-relevance">Speed, privacy, and relevance</h3><p>Beyond being able to find pizza on a Saturday night, we believe everyone should be able to browse the Internet freely – and not have to sacrifice the speed, privacy, or relevance of their search results in order to do so.</p><p>In the near future, we’ll be investing in features to bring even more of the benefits of Cloudflare infrastructure to every 1.1.1.1 + WARP user. Stay tuned!</p>]]></content:encoded></item></channel></rss>