Program
Note: Papers will be made available with the proceedings on September 25th.
You have not added any papers yet.
Under View Full Program, tick the checkbox next to each paper you plan to attend to create your custom program. Selections are saved in your browser's localStorage.
- 09:00 - 17:00 - IMC 2026 Workshops
- 18:00 - 20:00 - Welcome Reception
- 9:00 - 9:20 - Opening Remarks
- 9:20 - 10:30 - Keynote
- 10:30 - 11:00 - Break
- 11:00 - 12:30 - Session 1
- Session 1: Internet Topology and the AS Ecosystem (Session Chair: TBA)
-
LongCaleb Wang (Northwestern University), Ying Zhang (NORTHWESTERN UNIVERSITY), Qianli Dong (Northwestern University), Esteban Carisimo (Northwestern University), Ramakrishnan Durairajan (University of Oregon), Fabian Bustamante (Northwestern University)Abstract: As the backbone of global Internet connectivity, the submarine cable network (SCN) faces growing threats with serious economic and security implications. Strengthening its resilience requires a clear understanding of which cables and landing points are most critical. This depends on accurately mapping traffic onto the underlying infrastructure to identify vulnerabilities and assess their regional and global impact. Yet existing methods often lack the resolution and reliability needed for such analysis, leaving researchers and policymakers without the insights to safeguard this vital system. This paper introduces Calypso, a framework for mapping traceroute paths to the submarine cables they traverse. Calypso integrates ownership records, routing metadata, and geographic constraints to infer cable usage despite the opacity of the SCN and challenges such as route virtualization and inland infrastructure. It also defines Route Stress, a traceroute-derived metric for estimating the relative importance of submarine cables. Through expert validation, failure analysis, and regional case studies, we demonstrate Calypso's utility in revealing SCN dependencies and informing resilience efforts.
-
LongYiluo Wei (The Hong Kong University of Science & Technology (Guangzhou)), Jiahui HE (The Hong Kong University of Science & Technology (Guangzhou)), Amreesh Phokeer (Internet Society), Theophilus A. Benson (Carnegie Mellon University), Gareth Tyson (The Hong Kong University of Science & Technology (Guangzhou))Abstract: Africa's rapid development underscores the critical need for high-quality connectivity to support education, innovation, and economic opportunities. Despite substantial investments in Internet infrastructure, challenges persist, driven by a limited understanding of how particular infrastructural advancement strategies impact Internet-layer metrics. We present the first large-scale longitudinal measurement study to assess the effectiveness of recent (2018–-2023) Internet infrastructure deployments across Africa. We analyze long-term traceroute data, alongside DNS measurement data, to evaluate improvements in latency, routing, and local content resolution. Our study covers the deployment of 9 new submarine cables involving 19 countries, and 2058 new AS peering arrangements across 69 Internet Exchange Points (IXPs) in 32 countries. We provide critical insights into the evolving African Internet landscape, revealing the extent to which new infrastructure enhances connectivity and performance. These insights assist in guiding future investment decisions and ensuring sustainable improvements in Africa's Internet ecosystem.
-
LongZhiyi Chen (Georgia Tech), Zachary Bischof (Georgia Tech), Cecilia Testart (Georgia Tech), Alberto Dainotti (Georgia Tech)Abstract: Autonomous Systems (ASes) form the backbone of the Internet's routing infrastructure and aggregate diverse network services. In this work, we rethink how network properties are attributed to ASes in light of the broad range of Internet datasets now available, and argue that each property of interest should be represented as its own tag rather than as part of a rigid taxonomy. Guided by this view, we provide principles and a shared, unified data-and-software framework that facilitates systematic, accurate, flexible, and reproducible AS classification by network properties. Empirical analyses and multiple case studies demonstrate its research impact and improvements over the state of the art.
-
LongZhiyi Chen (Georgia Tech), Zachary Bischof (Georgia Tech), Cecilia Testart (Georgia Tech), Alberto Dainotti (Georgia Tech / Google)Abstract: Classifying Autonomous Systems (ASes) by the business sectors of their operating organizations, also referred to as AS business classification, is a relatively recent but increasingly important area of study. However, the only publicly available dataset for this purpose suffers from significant inaccuracies and limited maintainability due to its heavy reliance on third-party sources. We argue that an organization’s web presence—particularly its official website—serves as a more reliable and transparent source for business classification, offering broad coverage and regularly updated information. We introduce AS2Biz, a web presence-based methodology for AS business classification. AS2Biz first constructs an AS-to-website mapping dataset, achieving 98.5% coverage and 94.4% accuracy, and then mainly leverages large language models to classify textual website content into business categories. The entire methodology is highly automated and cost-efficient. Our final AS2Biz dataset covers 94.3% of ASes and, on a 220-AS ground truth benchmark, achieves substantially higher accuracy—more than doubling F1 score on subdivided business categories (0.78 vs 0.37) and improving F1 score on general business sectors (0.87 vs 0.62).
-
LongEric Pauley (Virginia Tech & Terrace Networks), Yohan Beugin (University of Wisconsin-Madison), Paul Barford (University of Wisconsin-Madison), Patrick Traynor (University of Florida), Patrick McDaniel (University of Wisconsin-Madison)Abstract: In 2021, security researchers disclosed a broad class of vulnerabilities in the management of configurations to major cloud providers. Because cloud providers reuse public IP addresses across tenants, and those tenants create configurations that associate trust with these addresses, adversaries could allocate IPs and abuse these latent configurations in myriad ways. However, it is unclear whether providers and customers have improved security practices in response to these vulnerabilities, and whether discovered vulnerabilities are exploited in practice. In this work, we perform a three-year longitudinal study of the defender and attacker response to latent configuration vulnerabilities. We leverage data from the DScope Internet telescope and certificate transparency logs, and introduce honey configurations, a new active scanning approach that directly measures exploitation of latent configurations. We find that cloud providers and customers have failed to respond to the threat of IP address reuse (with 60% of originally-observed vulnerabilities present years after disclosure). Meanwhile, adversaries are seeking out and exploiting IP address reuse, with a 39x increase in exploitation since prior disclosures. Our active measurement study finds evidence of global, multi-cloud adversarial apparatus designed to identify and exploit latent configuration. Attackers host malicious content, mal-issue TLS certificates, and execute cookie-based attacks on target websites. Our findings demonstrate the practical impact of these vulnerabilities, and we urge security practitioners to take steps to protect cloud deployments.
-
Abstract: While the Internet is designed to be a global network, peering disputes within its core infrastructure can undermine global reachability. The prolonged dispute between Cogent and Hurricane Electric serves as a notable example for fragmenting the IPv6 topology, although the precise impact on reachability has not been quantitatively assessed. In this paper we study the integrity of the Internet by measuring the level of fragmentation in the IPv6 Internet topology and estimating its impact on end-user services. Using BGP data we measure reachability differences between large transit providers in both IPv4 and IPv6. In addition we leverage forward DNS data to estimate the impact the IPv6 fragmentation can have on popular website reachability. We found that IPv6 fragmentation is substantial compared to the more unified IPv4 infrastructure. Furthermore, this fragmentation is worsening over time, with an increasing number of popular websites becoming partially reachable. Although current fallback mechanisms to IPv4 mitigate this issue, this study documents a fundamental challenge for a fully functional IPv6 Internet.
- 12:30 - 14:00 - Lunch
- 14:00 - 15:30 - Session 2
- Session 2: Policy and Society (Session Chair: TBA)
-
LongEmmanouil Papadogiannakis (FORTH and University of Crete), Panagiotis Papadopoulos (FORTH-ICS), Nicolas Kourtellis (Keysight AI Labs), Evangelos Markatos (FORTH and University of Crete)Abstract: Advertorials represent a marketing strategy where advertisements are designed to resemble the style and tone of editorial content. Despite their appearance, they are, in fact, paid content intended to promote a product, brand or service. Studies indicate that advertorials are more effective (81%) and less intrusive than traditional banner ads or pop-ups. Despite regulatory efforts for clear disclosure of paid content, concerns persist about the deceptive nature of advertorials. Advertorials can mislead readers into believing that they are consuming unbiased editorial content. In doing so, they gain undeserved legitimacy, by draping themselves in the credibility of the publication's design. In this study, we conduct the first systematic large-scale study of advertorials. We propose a novel automated methodology for detecting problematic advertorials in the wild, and collect 186K ad URLs over a period of 5 months. We investigate their prevalence and explore their structural and linguistic characteristics, finding that advertorials appear in 1 out of 3 news websites, including popular and credible outlets (e.g., The Guardian, EuroNews, CNN). Additionally, they often exhibit behavior associated with malicious practices and deliberately obscure or make legal disclaimers difficult to recognize, preventing users from identifying the promotional nature of the content.
-
LongYe Shu (UC San Diego), Elisa Luo (UC San Diego), Paul Chung (UC San Diego), Geoffrey M. Voelker (UC San Diego), Stefan Savage (UC San Diego)Abstract: In the U.S., the public right of access to court proceedings arises from the common law and the First and Sixth Amendments. However, courts have long recognized limited exceptions to this right, and certain court cases may be “sealed” (inaccessible to all but authorized parties) and thus, even the existence of such proceedings may be shielded from public scrutiny. This paper develops empirical techniques to infer both the prevalence of Federal case sealing and the duration of such sealing orders in practice. Using a combination of comprehensive crawling, undocumented APIs, and heuristics for inferring per-court case-numbering policies, we are able to use existing online legal databases (the U.S. Courts’ CM/ECF and IDB systems and the Free Law Project’s CourtListener RECAP service) to tease out the presence of case sealing via “absences” in the online record. Using this approach, we provide an initial analysis of sealing practices across twenty Federal court districts, how these relate to different “kinds” of proceedings, and how our contemporary findings compare to the U.S. Courts’ internal study conducted seventeen years prior.
-
LongJürgen Brandl (TU Wien), Florian Holzbauer (IT:U), Aaron Kaplan (DIGIT S.2 / European Commission), Johanna Ullrich (IT:U), Adrian Dabrowski (University of Applied Sciences Technikum Wien)Abstract: The European Union aims for digital sovereignty to reduce geopolitical risk, but anecdotal evidence suggests a significant gap between regulation and implementation. Yet, its precise extent remains undetermined due to the challenging, multifaceted nature of external dependencies. We present the first systematic assessment of digital sovereignty, empirically analyzing 1,988 domains across four key sectors (government, banking, media, higher education) in all 27 EU member states and EU institutions. We collect DNS records, TLS certificates, and HTTP responses to analyze six sovereignty dimensions: nameservers, web hosting, external resources, certificate authorities, email, and SaaS. Notably, we identify internal third-party SaaS deployments via domain verification tokens, offering insight into organizations’ IT landscapes. Dependencies are classified by criticality and EU/non-EU status to assess sovereignty from the infrastructure to the application layer. Our results show a heavy reliance on non-EU providers, with only 61 (3%) domains having no detectable dependencies. Key exposures include CDN/DDoS services, cloud-hosted email, office suites, and TLS certificates, all of which risk sudden interruption in volatile geopolitical times. This exposes a systemic sovereignty gap complicating NIS2 and DORA compliance. As a possible remediation, we elaborate on service agility and provide an open dataset for replication, monitoring, and informed policy decisions.
-
LongOrlando E. Martínez-Durive (IMDEA Networks Institute), Iñaki Ucar (Universidad Carlos III de Madrid), Zbigniew Smoreda (Orange Research), Esteban Moro (Northeastern University), Marco Fiore (IMDEA Networks Institute)Abstract: Understanding how digital platform consumption relates to political alignment is paramount in contemporary democracies, yet large-scale observational evidence spanning multiple platforms and elections remains scarce. We analyze passive mobile network metadata from a major French mobile network operator (MNO) covering approximately 31% of the mobile market across urban and suburban areas of metropolitan France during the 2019 and 2024 European Parliament elections. Integrating demands for tens of mobile services, including social media, news, messaging, and streaming, with socioeconomic indicators, we model vote shares via Dirichlet regression. We find that digital consumption provides independent and complementary signals of political alignment that, for several parties, match or exceed the predictive power of traditional socioeconomic indicators; combining both increases explanatory ability by up to 28.23%. For instance, right-wing support is associated with higher consumption of Facebook and TikTok, whereas engagement with online news outlets, Twitter, and Instagram correlates with centrist and progressive vote. These associations are consistent across 2019 and 2024, providing a population-scale view of the relationship between mobile platform use and political alignment.
-
Abstract: Due to U.S. sanctions and strict internet censorship, Iranian iOS users are barred from accessing the Apple App Store and developer services. In response, despite violating Apple’s developer terms, a thriving underground ecosystem of third-party iOS app stores has emerged to serve Iranian users. This paper presents the first comprehensive empirical study of these clandestine app stores. We document how these stores operate, including their distribution mechanisms, user authentication processes, and evasion techniques. By collecting and analyzing more than 1700 iOS application packages and their metadata from three major Iranian third-party app stores, we characterize the ecosystem’s size, structure, and content. Our analysis reveals a significant presence of Iranian-exclusive apps, widespread distribution of cracked apps, unauthorized monetization of paid content, and embedded third-party tracking and piracy libraries. We also uncover a notable overlap among financial, navigational, and social apps that exist solely in this ecosystem, reflecting the unique digital constraints of Iranian users. Finally, we quantify the potential revenue losses for developers due to piracy and document security and privacy risks associated with altered binaries. Our findings highlight how sanctions, censorship, and enforcement gaps have enabled a parallel app distribution ecosystem with complex socio-technical implications.
-
ShortTarek Ramadan (Concordia University), AbdelRahman Abdou (Carleton University), Mohammad Mannan (Concordia University), Amr Youssef (Concordia University)Abstract: Platforms for security analysis, archiving, and paste sharing—such as VirusTotal, URLScan.io, Hybrid Analysis, the Wayback Machine, and RedHunt—can unintentionally expose links containing sensitive information, yet no large-scale measurement has quantified the extent of such exposures. We present an automated system that detects and analyzes potential sensitive information leaked through publicly accessible URLs. The system combines lexical URL filtering, dynamic rendering, OCR-based extraction, and content classification to identify potential leaks. We apply it to 6,094,475 URLs collected from public scanning platforms, paste sites, and web archives, identifying 12,331 potential exposures across authentication, financial, personal, and document-related domains. These findings show that sensitive information remains exposed, underscoring the importance of automated detection to identify accidental leaks.
- 15:30 - 16:00 - Break
- 16:00 - 17:00 - Session 3
- Session 3: Internet-Wide Scanning and Network Telescopes (Session Chair: TBA)
-
LongChase Kanipe (Johns Hopkins University), Erik Rye (Johns Hopkins University), Dave Levin (University of Maryland), Robert Beverly (San Diego State University)Abstract: The vast, sparsely populated, and often ephemeral IPv6 address space makes discovering active addresses challenging. In response, the community has developed over thirty different IPv6 Target Generation Algorithms (TGAs). TGAs learn addressing structure from seed (training) datasets, build a representative model, and generate candidate addresses. Unfortunately, the existing literature employs a wide variety of input seeds, data cleansing, and metrics of ``success'' that prevent ready comparison. Toward facilitating unbiased algorithm performance evaluation, we present 6SEVEN, an extensible framework that hosts TGAs as plugins and links them to a shared framework with data cleaning, probing, dealiasing, and result tabulation components. We port as 6SEVEN plugins eight popular TGAs and perform a controlled comparison. Our results show that the performance of these TGAs varies drastically due to factors independent of the main algorithm -- especially the seed set composition, dealiasing, and tabulation procedures. We identify fundamental tradeoffs between desirable features of TGA performance, including yield and novelty. Our results call into question any simple ranking of TGAs and instead suggest that different TGAs are better suited to different use cases.
-
ShortLion Steger (Technical University of Munich (TUM)), Christian Junginger (Technical University of Munich (TUM)), Georg Carle (Technical University of Munich), Johannes Zirngibl (Max Planck Institute for Informatics)Abstract: Scanning the IPv6 address space exhaustively is infeasible. Besides hitlists with known active addresses, Target Generation Algorithms (TGAs) are a common tool for generating input for IPv6 measurements. They generate candidate addresses based on known responsive addresses using address patterns or machine learning. Some TGAs additionally use active scans during address generation as feedback mechanisms. However, the Internet is a changing ecosystem and scans can be impacted by network events, e.g., rate limits. These effects impact TGA results, a proper evaluation of their quality and benefit, and the comparability of TGAs. We propose an IPv6 Punching Bag, a local environment to test IPv6 scans and TGAs before actually using them on the Internet. It allows configuring prefixes with different response rates, e.g., aliased prefixes, or address patterns. It can be used on a single machine with a low memory footprint. We evaluate six dynamic TGAs with the Punching Bag and show that their adherence to scanning budgets and limits, and their inability to detect aliased prefixes necessitates local tests before Internet scans. We also demonstrate that 6Scan, adynamic TGA, is not adapting to different response behavior.
-
LongEric Kapitanski (University of Southern California- Information Sciences Institute), Jelena Mirkovic (USC Information Sciences Institute)Abstract: Traffic to unused IP addresses (dark spaces), including passive and reactive network telescopes and honeynets, has traditionally been used to study malicious activities. However, attackers may avoid known dark space sensors and preferentially target endpoints hosting legitimate services. We present LightScope, a lightweight application that unifies passive observation at Layer 4 with selective proxying to a remote Layer 7 honeypot. LightScope does this by leveraging a subset of the unused ports on production servers. By embedding measurement on the production hosts themselves, LightScope creates a new vantage point for observing various scanner populations, port-targeting profiles, and post-login behaviors that are only partially visible from dark-space telescopes and honeynets. In this paper we analyze two months of data collected from 235 LightScope deployments. We show that LightScope contributes new knowledge about widely observed scanners, their port scan strategies and their post-exploit activities. These scanners interact with production endpoints running LightScope differently than they do with other vantage points. In our study period, LightScope also exclusively observes 64,686 scanners not present at other vantage points. By instrumenting production machines, LightScope provides complementary visibility to existing telescopes and honeynets.
-
ShortDario Ferrero (Delft University of Technology), Andrea Sordello (Politecnico di Torino), Harm Griffioen (Delft University of Technology), Idilio Drago (Università di Torino), Georgios Smaragdakis (Delft University of Technology), Marco Mellia (Politecnico di Torino)Abstract: Reactive telescopes have characterised TCP scanning at scale by replying to incoming SYN packets. The UDP equivalent remains unexplored, and, in general, little is known about UDP scanners. We present the first transport-layer reactive UDP telescope, RT-UDP, which we deploy alongside pure darknets and TCP responders at four networks in three countries. By means of measurements over 30 consecutive days, we compare traffic seen by pure and reactive telescopes. Our contributions are threefold. Firstly, we measure the activation effect. Scanning peaks at up to $3\times$ the darknet baseline; sustained activity concentrates on fewer ports and surfaces services largely invisible to passive telescopes (e.g., Ethereum discovery traffic rises from negligible levels to millions of packets). Secondly, we characterise the sources. While well-known scanners account for 30--40\% of daily UDP traffic, RT-UDP isolates a separate population producing amplification bursts. Thirdly, we observe behaviours invisible to passive telescopes. Many protocols are probed across 10--100$\times$ more ports than in the darknet, and some scanners deploy multi-stage probes that switch protocol after a response or use innocuous packets before sending malicious follow-ups.
- 18:00 - 20:00 - Conference Dinner
- 9:00 - 10:30 - Session 4
- Session 4: TLS, PKI, and Protocol Authentication (Session Chair: TBA)
-
LongNimesha Wickramasinghe (The University of New South Wales (UNSW)), Frank Li (Georgia Institute of Technology), Sanjay Jha (The University of New South Wales (UNSW)), Arash Shaghaghi (The University of New South Wales (UNSW))Abstract: Post-quantum cryptography (PQC) has evolved from a long-term planning concern into an operational priority. Following NIST’s standardization of PQC algorithms, governments and standards bodies published transition roadmaps outlining migration timelines, priority sectors, and deployment strategies. However, our survey of these policies reveals substantial divergence in technical prescriptions and urgency. It remains unclear how widely PQC has been adopted in practice and how policy differences translate into observable deployment outcomes. To quantify trends in public-facing post-quantum TLS (PQ-TLS) adoption, we conduct three measurement rounds between July 2025 and March 2026, covering one million public HTTPS endpoints. Across these rounds, we establish more than two billion TLS 1.3 handshakes from 11 globally distributed vantage points to characterize cryptographic negotiation behavior. We observe strong convergence in hybrid key exchange, with every observed post-quantum negotiation selecting X25519MLKEM768, but find no verified deployment of post-quantum signature algorithms for authentication. Hybrid key-exchange adoption is highly concentrated among endpoints attributed to managed infrastructure providers. Country and sector comparisons show limited correspondence between early deployment patterns and published transition timelines. Contrary to early experimental studies suggesting measurable overhead, our experiments show no meaningful latency increase for hybrid key exchange in Internet settings. We further observe that PQC adoption frequently coexists with legacy TLS configurations. Together, these findings highlight a gap between policy expectations and early deployment reality, and provide empirical insight to inform more grounded PQ-TLS transition.
-
ShortJonas Mücke (TU Dresden), Konstantin Gasser (TU Dresden), Thomas C. Schmidt (HAW Hamburg), Matthias Wählisch (TU Dresden)Abstract: We present a new measurement approach to detect ECH deployments. This method leverages standard-compliant behavior of ECH servers. Prior measurements relied on DNS exposure and reported only one major ECH deployment (Cloudflare). Our measurements reveal yet undiscovered ECH deployments, including a large deployment at Meta. Meta servers support ECH, but ECH configuration is not exposed via the DNS. We also study deployment details such as privacy improvement and latency penalties due to~ECH.
-
Abstract: DNS-based Authentication of Named Entities (DANE) offers an alternative or complementary authentication mechanism. However, DANE’s web adoption is limited. In this work, we revisit DANE by focusing on a practical deployment concern for the web: added latency. Using our custom DANE-for-the-web measurement framework, we show that DANE’s overhead is dominated by DNS resolution, particularly DNSSEC, which is highly sensitive to network latency and caching behavior. We further reveal that certain TLSA parameter combinations can trigger DNS-over-TCP fallback due to large response sizes, increasing DNS delay by up to 2 round trip times. Finally, our web-scale study discovers 4,435 websites supporting DANE validation for HTTPS connections, and characterizes DANE-induced delays across these real-world domains. Our findings shed light on the readiness of DANE for the web and offer guidance for efficient deployment.
-
ShortNehal Fooda (University of Maryland), James Larisch (Cloudflare, Inc.), John Schanck (Mozilla), Bruce Maggs (Duke University and Emerald Innovations), Tijay Chung (Virginia Tech), Dave Levin (University of Maryland)Abstract: Certificate revocation is the last line of defense against key compromise. When a website administrator instructs their Certificate Authority (CA) to revoke a certificate, the CA must make the revocation data information publicly available to clients through distribution mechanisms like CRLs and OCSP. Not only is it important that the website, CA, and client all do what is necessary to produce and consume revocation state, it is also important that these all be done in a timely manner. Extensive measurements have been made on how quickly (and whether) websites revoke compromised certificates, whether clients check for revocations, and even the reachability of OCSP servers. But to our knowledge, there have been no studies on how quickly CAs make turn revocation requests into publicly downloadable revocation data. In this paper, we perform the first study of the latency from a website’s revocation request to when the CA makes it available to clients. We develop a custom tool for revoking and downloading revocation state, and apply it to six of the most popular CAs on the web today across both CRL and OCSP. We measure whether CAs delays in making revocation data available is a function of time of day, distribution mechanism, reason code, and other factors. We also measure the end-to-end latency from when a website requests their certificate to be revoked to when Mozilla Firefox’s CRLite revocation checking system obtains it. Collectively, we find that CAs operate on tighter time bounds than currently strictly required by the CA/B Baseline Requirements, but that some major CAs are still leaving very large (up to 24-hour) delays from revocation request to revocation availability.
-
LongMuhammad Hamza (Virginia Tech), Mattijs Jonker (University of Twente), Raffaele Sommese (University of Twente), Simon Fernandez (Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG), Olivier Hureau (Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG), Eric Pauley (Virginia Tech & Terrace Networks), Taejoong Chung (Virginia Tech)Abstract: Email remains the backbone of online communication, yet its authentication mechanisms continue to lag behind evolving attacker capabilities. Email authentication protocols are layered, with their interactions propagating and amplifying underlying vulnerabilities. This is clearly evident in the case of SPF vulnerabilities, where pairing them with DMARC's organizational context of authentication can lead to organization-wide spoofing. In this paper, we perform a comprehensive empirical analysis of this phenomenon. Despite SPF vulnerabilities being well known, our large-scale dataset reveals 528~k vulnerable policies. When paired with DMARC, these policies enable spoofing of 79~k domain hierarchies. To mitigate this, we propose Split Alignment, a practical configurational technique that protects established email-sending domains from all SPF failures in the domain hierarchy. We then use our novel methodology to validate the security and deliverability of Split Alignment in a consistent and reproducible manner. In doing so, we identify and disclose critical bugs in the DMARC authentication flows of 4 major email services. In this way, we strengthen the security of the email ecosystem by mitigating long-standing SPF vulnerabilities.
-
LongHamim Hamid (North Carolina State University), David Adei (North Carolina State University), Bradley Reaves (North Carolina State University)Abstract: In response to an epidemic of unlawful unsolicited telephone calls, the United States mandated the use of a novel, previously unimplemented call attestation framework called STIR/SHAKEN (S/S) in 2019. The mandate required telephone network operators to deploy a brand-new, federated, industry-wide PKI within timelines shorter than 2 years for large providers. Under those circumstances, achieving any degree of successful deployment would be remarkable. In this paper, we present a longitudinal study of S/S deployment, evaluating standards compliance, certificate authority practices, and provider call signatures and metadata. Our primary data source is signaling data from 7.75 million calls collected by multiple telephone honeypots over four years (November 2021–February 2026). Our results show rapid adoption of S/S, but with pervasive non-compliance with standards. We also found repeated instances of problematic practices, including providers who sign calls with root keys or originate calls with pre-signed call attestations from unrelated calls. Despite these lapses, we find that most calls are correctly signed and that certificate authority and operator practices and standards compliance are improving at a modest pace. Our discussions highlight concrete directions to strengthen policy and deployment, enabling more effective mitigation of telephone network abuse.
- 10:30 - 11:00 - Break
- 11:00 - 12:30 - Session 5
- Session 5: Web Measurement and Performance (Session Chair: TBA)
-
LongRumaisa Habib (Stanford University), Balaji Balachandran (Stanford University), Nurullah Demir (Stanford University), Zakir Durumeric (Stanford University)Abstract: Modern websites load hundreds of resources, including images, JavaScript, stylesheets, and fonts. While these web resources enable rich interactive experiences, they also carry heavy network, memory, and CPU costs. Many use cases---like content archival and programmatic automation---do not need the full intended functionality of the website to be successful. Yet these unnecessary resources are fetched regardless, and at scale, this over-fetching compounds into a substantial environmental footprint. However, it is unclear which resources are necessary for which use cases. In this paper, we provide the first modality-aware measurement of unnecessary resources on the web. To do so, we introduce four dimensions of webpage fidelity to characterize modalities on the web. We show that we can reduce 17-98% of network bytes, 39-76% of CPU time, and 11-53% of memory usage while maintaining the relevant content for a given modality. We estimate that modality-aware selective loading could avoid up to 159 million metric tons of CO2e annually. These findings uncover additional web bloat beyond what was previously understood, and opportunities to build more efficient and sustainable tooling for crawling, interacting with, and archiving content on the web.
-
LongAdrian Kunz (University of Kassel), Lukas Kawollek (University of Kassel), Yasin Alhamwy (University of Kassel), Oliver Hohlfeld (University of Kassel)Abstract: Over the last 15 years, the web has changed significantly with the use of mobile-first design, JavaScript-heavy single-page applications, HTTP/2 and HTTP/3, and widespread CDN-driven infrastructure. In this replication track paper, we provide a fresh perspective on website complexity based on the current web and revisit the widely cited IMC’11 paper by Butkiewicz et al . [14 ]. We find non-origin resources becoming increasingly dominant since 2011, now exceeding same-origin content in importance and contributing a substantial share of total traffic, with JavaScript alone accounting for 77.9% of non-origin bytes and Tag Managers appearing on 67.1% of websites. We also analyze which metrics are most critical for predicting page load behavior and find that the number of objects is still the most important factor, followed by total bytes and origins, although correlations have weakened over time. We find that subpages are generally less complex than landing pages, while client-side filtering can significantly reduce the complexity, and mobile and desktop versions show similar complexity. Modern infrastructure is widely adopted, with HTTP/2 used in 58.8% and HTTP/3 in 32.7% of requests, while minification still yields limited benefits for many sites. Modern frontend frameworks are used by around 30% of websites, jQuery remains dominant at 66.8%, and most JavaScript-heavy pages (77.2%) grow after execution, with a median increase of 31.9%.
-
LongKarthik Ramakrishnan (Georgia Institute of Technology), Qinge Xie (Georgia Institute of Technology and Zhejiang University), Uma Anand (Georgia Institute of Technology), Frank Li (Georgia Institute of Technology)Abstract: Web privacy measurements fundamentally depend on web crawling techniques. Unlike many other web measurements that can be performed only using the landing page (e.g., TLS assessments, web server security configurations), web privacy evaluations often must explore deep into a website to more comprehensively identify and characterize its privacy behaviors. However, prior studies adopt various crawling configurations and strategies, and to date, we lack a systematic understanding of the impact of different crawling configurations on the privacy behaviors observed. In this paper, we address this gap by conducting controlled web crawling experiments across the top 10K CrUX websites, while monitoring privacy-related metrics including HTTP request domains, JavaScript Web APIs accessed, and cookies set. We evaluate 12 different crawling strategies, varying the page search strategy (depth-first, breadth-first, and a hybrid approach), whether browser state is preserved across page visits, and whether user interaction is simulated. We further assess the impact of varying the number of pages crawled per site, as well as the amount of time the crawler waits per page. Our analysis shows how different crawling parameters influence different privacy metrics, ultimately informing recommendations for designing web crawling methods for web privacy measurements.
-
LongMeher Afroz (Stony Brook University), Muhammad Daniyal Pirwani Dar (Stony Brook University), Omar Chowdhury (Stony Brook University)Abstract: Web accessibility has been a focal point of numerous discussions, leading to several key compliance requirements that developers of websites must respect. Prior work, including WebAIM's yearly reports, summarizes noncompliance with web accessibility requirements using the arithmetic mean and concludes that violations of web accessibility requirements continue to rise. We analyze 100K website homepages from 1996 to 2025 using historical snapshots spanning multiple categories and regions from the Wayback Machine. We observe that the underlying distributions of accessibility violation/error counts in homepages of websites are strongly right-skewed and follow a discrete log-normal distribution, under which the mean, median, and geometric mean summarize different aspects of the data. These statistics support different conclusions on this corpus, since the mean rises while the geometric mean and median do not. The reported direction of change is therefore also a function of the summary statistic, not of the data alone. Using the geometric mean, we observe that the trend of the total number of accessibility errors flattens in recent years. This aggregate, however, conceals a divergence across the four WCAG principles: Perceivable category violations fall while Operable category violations roughly triple between 2000 and 2015. The total number of accessibility violations alone, therefore, does not establish web-accessibility improvement or deterioration and requires a finer-grained analysis. We also observe that existing web accessibility error checkers disagree with one another, where some success criteria are detected by only a single tool, and the best available tool detects only 24.8% of possible violations. Overall, although there are signs of progress, a fully inclusive web, unfortunately, still remains out of reach.
-
LongHyunSeok (Daniel) Jang (NORTHWESTERN UNIVERSITY), Ayush Pandey (New York University Abu Dhabi (NYUAD)), Matteo Varvello (Nokia), Fabi'an E. Bustamante (Northwestern University), Yasir Zaki (New York University Abu Dhabi (NYUAD))Abstract: Browser extensions are widely used, but developers still face limited options for monetization. Mellowtel recently introduced a new approach: \textit{bandwidth-as-a-service}. By embedding its SDK, developers earn revenue from opted-in users who share idle bandwidth to fetch web data for AI labs and startups. We combine large-scale Chrome Web Store measurements, controlled experiments across six AWS regions, and an in-the-wild deployment with 34 students spanning 25 countries to reveal how the SDK operates and the extent of its adoption. We find that while hundreds of extensions bundle the Mellowtel SDK, active monetization is rare: only 16 of 116 Mellowtel-equipped extensions actively inject crawl iframes during a browsing session, collectively reaching far fewer users than the millions Mellowtel claims, making extension counts an upper bound on real participation. Crawling intensity varies sharply by geography, with US-based nodes receiving up to 8.7$\times$ more crawl tasks than other regions. Mellowtel's crawl assignments split into two distinct workloads: broad and shallow scraping of a long tail of domains, and repeated structured extraction from a small set of high-demand targets (e.g., Google search, CouponFollow, Reddit). Finally, we identify concrete security concerns: crawl target lists include domains flagged as associated with known malware families, and some users received crawl requests for content banned in their countries. More critically, a dynamic code loading mechanism allows Mellowtel to execute remotely fetched JavaScript inside the user's browser at runtime, bypassing Chrome Web Store review entirely.
-
ShortAyush Pandey (New York University Abu Dhabi (NYUAD)), Vladimir Sharkovski (New York University Abu Dhabi (NYUAD)), Matteo Varvello (Nokia), Yasir Zaki (New York University Abu Dhabi (NYUAD))Abstract: JavaScript (JS) bloat is a major source of slow, data-inefficient web performance, as many sites ship large monolithic bundles with substantial unused code. Existing JS dead-code elimination techniques help but require server-side support and are therefore hard to deploy. This paper evaluates HTTP range requests as a client-side alternative to JS dead code elimination, without requiring server changes beyond standard protocol support. We analyze over 421,000 domains serving JS in the top one million websites and find that roughly two-thirds support even complex multi-range HTTP requests. On a 5,000-page subset analyzed with Google Lighthouse, we estimate that fetching only essential code could reduce JS transfer by about 20%. We then implement ThorJS, a fully client-side technique for JS dead-code removal. It intercepts JS requests and attaches range headers to retrieve only essential code segments of JS resources. Using this prototype, we evaluate system overheads in the browser stack, and application-level performance on 100 representative webpages in the wild. Results show that ThorJS reduces JS transfer by a median of 4 MB, with savings up to 38 MB, and improves page-load time by 13 seconds under slow 3G conditions, with negligible overheads for clients and servers.
- 12:30 - 13:30 - Lunch
- 13:30 - 14:30 - Community Session
- 14:30 - 15:30 - Session 6
- Session 6: Devices and Network Services: Security and Reliability (Session Chair: TBA)
-
LongAndrew Losty (University College London), Tianrui Hu (Northeastern University), Daniel J. Dubois (Northeastern University), Narmeen Shafqat (NUST), Aanjhan Ranganathan (Northeastern University), David Choffnes (Northeastern University), Anna Maria Mandalari (University College London)Abstract: The rapid growth of the Internet of Things (IoT) has created fragmented smart home ecosystems, with platforms such as Apple Home, Google Home, and Samsung SmartThings historically relying on proprietary protocols and device-specific implementations that limited interoperability. The Matter standard introduces a common, secure framework enabling cross-platform compatibility, now supported by over 600 manufacturers. This study empirically analyzes Matter and non-Matter ecosystems, focusing on security and privacy. We evaluate 25 Matter devices and 14 non-Matter devices with comparable functionality, examining network traffic during commissioning, device control, and data transmission. Our evaluation shows that Matter improves security and privacy through AES-128-CCM encryption, device attestation, and secure commissioning. Its local-first design further reduces cloud dependency and limits data exposure. However, we identify several risks. Persistent multicast DNS broadcasts reveal detailed device metadata, and rotating key identifiers are inconsistently implemented, with some devices failing to renew them for extended periods. We also observed ecosystem-specific differences in commissioning behavior, and many devices maintain vendor or third-party cloud connections. Firmware maintenance is inconsistent, as most devices operate on early Matter versions and did not receive updates during the study period. Moreover, encrypted unicast traffic patterns can disclose device types and user activities to an in-network adversary.
-
LongMartin Mladenov (Delft University of Technology), Max van der Horst (Delft University of Technology), Alexandru Hossu (Delft University of Technology), Rok Štular (Delft University of Technology), Harm Griffioen (Delft University of Technology), Georgios Smaragdakis (Delft University of Technology)Abstract: Full-stack web development frameworks have become increasingly popular in recent years. One such framework, Next.js, is extensively used by one in five web developers. In December 2025, a critical vulnerability with the highest severity score of 10 was discovered, affecting Next.js through supply-chain propagation. It allows attackers to remotely execute code on vulnerable servers without user authentication and with a single malicious HTTP request. Additionally, public proof-of-concept exploits contributed to widespread exploitation. In this paper, we investigate 13 million exploitation attempts targeting Next.js and characterize the malware deployed by attackers following successful exploitation. Leveraging our custom Next.js honeypots and reactive telescope, we identify 93 unique clusters of payloads performing a wide spread of exploitation activity such as botnet installation, cryptomining, and infostealers. Our analysis also shows that, initially, exploitation attempts relied on the released proof-of-concept code, but they later evolved into more sophisticated and effective code variants that successfully evade defense mechanisms such as Web Application Firewalls. We identify that some threat actors measure the vulnerability of targets before deploying any malicious payloads and demonstrate the presence of informed application-layer scanners with prior knowledge of a set of vulnerable hosts, which they exclusively target. Moreover, we observe markers consistent with LLM-style code artifacts, pointing towards the emergence of AI-assisted attacks. In the process, we characterize the demographics of the infrastructure used to attack Next.js and discuss the shortcomings of defense capabilities.
-
Abstract: Voice-over-IP (VoIP) powers nearly all modern telephony, yet its Internet-facing infrastructure lacks a systematic security census. We conduct the first Internet-wide measurement study of SIP security, scanning the IPv4 space to identify 631,045 application-layer responsive endpoints. Our analysis reveals a pervasive failure to adopt modern security standards: all but four authentication-requiring endpoints use 25-year-old MD5 defaults, with one SHA-256 and three SHA-512/256 responses, and 87% omit the qop replay defense. Furthermore, at least 76.6% of SIP-over-TLS endpoints also allow plaintext communication, and we estimate 1,808 endpoints with inferred open registration, some with strong PSTN connectivity indicators. We show that the landscape is highly centralized, with 10 providers controlling over 50% of the infrastructure, and demonstrate credential cracking and RTP eavesdropping against a controlled lab deployment.
-
ShortRuixuan Li (Tsinghua University), Enlong Li (Quan Cheng Laboratory), Mingxuan Liu (Zhongguancun Laboratory), Baojun Liu (Tsinghua University), Haixin Duan (Quan Cheng Laboratory; Tsinghua University), Qingfeng Pan (Coremail Technology Co. Ltd)Abstract: Reliable delivery is a basic expectation of email services. The increasingly complex ecosystem has led to more frequent email bounces. When encountering various errors, properly handling failures and retrying are crucial to ensuring email deliverability. Currently, the community lacks a comprehensive understanding of email retry strategies, and bridging this gap can enhance the reliability of email services. In this paper, we measure the email delivery retry strategies of 11 email service providers and 8 open-source email software. We deploy a platform (RetryWatch) that simulates various types of bounce errors returned by email receiving servers, covering the DNS, TCP, and SMTP communication stages. We find that many providers and software exhibit defects in their email retry implementations. For example, we find that six email systems exhibit poor load balancing across MX record entries in dual-stack environments. In addition, when encountering blocklistand greylist-related SMTP error messages, most email providers fail to handle retries appropriately. We release RetryWatch to support further testing by the email community. We hope this work helps improve email reliability and inspires efforts to refine email retry strategies.
- 15:30 - 16:00 - Break
- 16:00 - 17:30 - Session 7
- Session 7: Measuring AI Systems (Session Chair: TBA)
-
LongZilong Lin (University of Missouri-Kansas City), Zichuan Li (University of Illinois Urbana-Champaign), Xiaojing Liao (University of Illinois Urbana-Champaign), XiaoFeng Wang (Nanyang Technological University)Abstract: Open-source uncensored large language models (ULLMs) are increasingly available in public LLM ecosystems. By removing or weakening safety guardrails, these models can comply with harmful requests, generate unsafe content, and be readily reused in downstream systems. While prior work has demonstrated the misuse potential of individual uncensored models, little has been done to systematically understand the ULLM ecosystem in real-world settings. In this paper, we present a systematic study of the open-source ULLM ecosystem. We analyzed 22,096 ULLMs collected from five major LLM hosting platforms and examine them across five interconnected components: artifacts, development, propagation, governance, and adoption. Our analysis shows that ULLMs are produced through four major methods, supported by a rich ecosystem of publicly available tools and uncensored datasets. We further found that ULLMs propagate widely across platforms, forming 4,457 cross-platform clusters comprising 11,094 models, with Hugging Face serving as the dominant upstream source. We uncovered systemic governance and security gaps, including inconsistent moderation labels applied to ULLMs across platforms, cross-platform propagation of ULLMs flagged with threat indicators, and frequent violations of base-model licensing constraints. ULLMs are also actively deployed in diverse real-world settings, including web applications, GitHub projects, underground forums, and model routing platforms, significantly lowering the barrier to harmful use. Our findings indicate that ULLMs are not isolated unsafe models but components of a broader ecosystem in which supply-chain dynamics, weak governance, and diverse downstream deployments jointly amplify risks.
-
LongZuyao Xu (Nankai University), Xiang Li (Nankai University), Yuqi Qiu (Nankai University), Lu Sun (Nankai University)Abstract: Self-hosted large language model (LLM) serving is emerging as a distinct category of Internet service, but we still know little about how these deployments appear and change on the public Internet. We present a 365-day longitudinal measurement of exposed Ollama endpoints (port 11434) from February 2025 to February 2026, combining daily active probing with GeoIP/ASN enrichment, PTR and port-443 host observations, and survival analysis. Across 362 observation days and approximately 4.8~million IP$\times$day observations, 26.4\% of the 152{,}137 cumulative IPs appear for a single day; across five selected CVEs, only 0.43--2.90\% of below-fix IPs upgraded in place; the top five countries/regions account for over 70\% of weighted observations; and cloud and hosting providers dominate the top ASNs. These results characterize exposed Ollama as a structural exposure surface: persistent, growing, and heavily concentrated. At the same time, old versions, default model choices, cloud and hosting ASNs, PTR categories, and TLS certificate patterns remain visible across the year, indicating recurring insecure deployment practices in cloud infrastructure and the potential reach of provider-level mitigation.
-
LongHaofei Xu (Washington University in St. Louis), Umar Iqbal (Washington University in St. Louis), Jacob M. Montgomery (Washington University in St. Louis)Abstract: Google AI Overviews (AIOs) are arguably the most widely encountered deployment of generative AI, reaching over 1.5 billion users who may not realize the answers they see are AI-generated. Where search engines have traditionally surfaced ranked sources and left users to evaluate them, AIOs synthesize and deliver a single answer — giving Google unprecedented editorial control over what users read and know. We present a large-scale longitudinal measurement study, issuing 55,393 trending queries across 19 topical categories over a 40-day window (March 13–April 21, 2026). We report four main findings. First, overall AIO activation is 13.7%, rising to 64.7% for question-form queries, while politically sensitive topics see markedly lower rates. Second, AIO-cited domains are more credible than co-displayed first-page results, yet nearly 30% do not appear in those results at all, indicating a source selection mechanism distinct from Google’s ranking algorithm. Third, decomposing responses into 98,020 atomic claims, 11.0% are unsupported by the cited pages — with omission the dominant failure mode — and source quality and claim fidelity are largely independent. Fourth, well over half of AIO-cited pages carry display advertising, meaning publishers lose revenue when AIOs suppress the click-through, even as Google’s own sponsored ads continue to appear on the same page. Together, these findings document a rapid transformation of the online information ecosystem whose consequences for epistemic security remain poorly understood.
-
LongTingting Yao (George Mason University), Ruizhe Shi (George Mason University), Yao Liu (Rutgers University), Bo Han (George Mason University), Songqing Chen (George Mason University)Abstract: Smartglasses are emerging as hands-free and low-friction platforms for wearable voice AI assistants, but their end-to-end performance and system characteristics remain poorly understood. This paper presents the first empirical measurement study of five consumer-grade AI-enabled smartglasses--Ray-Ban Meta (RBM), Ray-Ban Meta Display (RBM DP), Demabon, EvenG1, and Halliday--and compares them with three phone-based counterparts: Meta AI, OpenAI ChatGPT, and Microsoft Copilot. We find that smartglasses consistently lag behind phones in responsiveness: their mean time-to-first-action (TTFA) spans 2.85--9.13 s, versus 1.79--2.79 s for phones. However, this gap is not a fixed wearable penalty, as smartglasses themselves exhibit large internal variation, from near phone-class TTFA at 2.85 s to slow-tier TTFA above 8 s. TTFA breakdown further shows that different devices are limited by different pipeline designs. For example, Demabon is server-bound, with 8.43 s server-side latency, whereas EvenG1 has a fast backend but is client-bound, with 4.28 s client-side latency. Under uplink packet loss, robustness depends not only on the transport protocol but also on query-submission granularity: QUIC-based RBM/RBM DP remain resilient, TCP-based EvenG1/Demabon degrade gracefully, while Halliday's single long-lived POST leads to cliff-like failure. Overall, our results call for smartglasses AI pipelines that co-design cloud services, cross-stage coordination, and robust transport.
-
ShortGuanjie Lin (University of Massachusetts Boston), Yinxin Wan (University of Massachusetts Boston), Shichao Pei (University of Massachusetts Boston), Ting Xu (University of Massachusetts Boston), Kuai Xu (Arizona State University), Guoliang Xue (Arizona State University)Abstract: Third-party Large Language Model (LLM) API gateways are rapidly emerging as unified access points to models offered by multiple vendors. However, the internal routing, caching, and billing policies of these gateways are largely undisclosed, leaving users with limited visibility into whether requests are served by the advertised models, whether responses remain faithful to upstream APIs, or whether invoices accurately reflect public pricing policies. To address this gap, we introduce VeriFlow, a lightweight black-box measurement framework for evaluating behavioral consistency and operational transparency in commercial LLM gateways. VeriFlow is designed to detect key misbehaviors, including model downgrading or switching, truncation, billing inaccuracies, and instability in latency by auditing gateways along four critical dimensions: response content analysis, multi-turn conversation performance, billing cccuracy, and latency characteristics. Our measurements across 10 real-world commercial LLM API gateways reveal frequent gaps between expected and actual behaviors, including silent model substitutions, degraded memory retention, deviations from announced pricing, and substantial variation in latency stability across platforms.
-
LongRuizhi Cheng (Meta Platform, Inc. and George Mason University), Guowu Xie (Meta Platform, Inc.), Bo Han (George Mason University)Abstract: Video-based conversational interaction between humans and generative artificial intelligence (GenAI) represents a rapidly emerging paradigm at the intersection of networking, multimedia communication, and vision language models (VLMs). Applications such as ChatGPT and Gemini enable users to point their smartphones at objects/scenes and receive contextual responses, opening new possibilities for on-the-go visual understanding. However, despite their growing adoption, the underlying performance dynamics and system behavior of these applications remain poorly understood. In this paper, we present the first systematic measurement study of video-based human–GenAI interaction across five representative applications with 11 distinct variants. We develop a custom toolset to characterize these applications across diverse network conditions and interaction scenarios, exposing how design decisions in video delivery and transport protocols shape user-perceived experience. Our measurements reveal several fundamental bottlenecks: conversational latency remains above natural turn-taking thresholds even in unconstrained networks; protocol choices significantly affect system robustness under packet loss; and conventional video streaming practices, such as prioritizing stable frame rates, often degrade VLM inference accuracy. Our findings provide an empirical foundation for designing and optimizing next-generation real-time human–AI interaction.
- 9:00 - 10:30 - Session 8
- Session 8: DNS in the Wild (Session Chair: TBA)
-
LongJoel Sommers (Colgate University), Calvin Kranig (University of Wisconsin-Madison), Wei-Shiang Wung (University of Wisconsin-Madison), Paul Barford (University of Wisconsin-Madison), Mark Crovella (Boston University), Eric Pauley (Virginia Tech & Terrace Networks)Abstract: DNS naming conventions reveal information about infrastructurecdeployments, operational configurations, and network topology. In this paper, we present \textit{namefill}, a tool that uses \textit{programming by example} to synthesize executable functions that summarize DNS naming conventions. Unlike regular expressions, which match but cannot generate, a synthesized namefill function produces the predicted fully-qualified domain name (FQDN) for addresses within the prefix it was synthesized from, enabling active probing for as-yet undiscovered FQDNs and systematic auditing of existing records. Applied to 1.2 billion PTR records, namefill synthesizes 313K distinct naming conventions across 254K registered domains. Using a PTR record dataset from CAIDA's ITDK, we find that namefill covers 67.7\% of IP/FQDN pairs and that its coverage is largely complementary to tools that focus on names that embed geotokens. We use the synthesized functions to systematically identify naming anomalies, including incorrect prefix bytes in PTR record names and inconsistencies in separator characters between IP address bytes. We identify 991 misconfigurations covering 236K PTR records and 1,639 inconsistencies covering 426K PTR records from 2026, and verify 43 misconfigurations with network operators. Lastly, we use the synthesized functions as a DNS record discovery tool to enumerate names in the DNS previously not captured in our source data. From 2.76M training names we find 194M previously unrecorded forward DNS names using 355M queries, a 70$\times$ increase.
-
LongZihan Li (Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences), Wenhao Wu (Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences), Zhaohua Wang (Computer Network Information Center, Chinese Academy of Sciences), Yu Tian (Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences), Chuan Gao (Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences), Qinxin Li (Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences), Yiming Xia (Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences), Zhenyu Li (Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences)Abstract: Modern DNS has evolved beyond traditional name resolution, increasingly serving as a carrier for structured metadata and service-discovery information. This functional enrichment inevitably drives sustained growth in DNS response payloads. While essential for modern ecosystem features, this growth is also making DNS response size the center of amplification risk. Despite its importance, the behavior of response size in the wild remains poorly understood, and we still lack a global, systematic view to guide its configuration and control. To this end, we conduct a year-long measurement of response-size behavior across about one million open DNS resolvers. We characterize the UDP response-size limits enforced by these resolvers and find that they cluster around a small set of common values. We also demonstrate how configuration changes at major public resolvers propagate across the DNS ecosystem. We further show that regional differences in resolver response-size limits lead to uneven amplification risk and expose fundamental limitations of existing static and upstream-based defenses. This complex response behavior stems from the tension between mitigating amplification risk through small fixed response-size limits and supporting diverse applications that require large responses. To address this tension, we propose Adaptive Response Scaling (ARS), a lightweight dynamic response-control mechanism that reduces amplification by over 70\% while preserving DNS functionality.
-
LongYizhe Zhang (University of Virginia), Hyeonmin Lee (Hanyang University), Shumon Huque (Salesforce), Yixin Sun (University of Virgina)Abstract: Authenticated denial of existence is a fundamental component of DNS Security Extensions (DNSSEC), realized through NSEC and NSEC3 records. Over time, several variants have emerged to improve the efficiency and privacy of denial-of-existence responses, including Compact Denial of Existence. Despite their operational importance, the prevalence and diversity of NSEC and NSEC3 deployments in the wild remain poorly understood. In this paper, we present a comprehensive measurement study of NSEC and NSEC3 deployment across 174 million domains in the .com, .net, and .org zones, as well as domains from four popular rankings. A key challenge in such measurement lies in accurately distinguishing among NSEC/NSEC3 variants across heterogeneous provider implementations. To address this challenge, we develop novel metrics driven by statistical observations, which reveal prevalent NSEC/NSEC3 variants such as Compact Denial of Existence in NSEC and White Lies in NSEC3. We also show that wildcard synthesis is non-negligible in the wild (potentially affecting nearly 19% and 15% of NSEC-enabled and NSEC3-enabled domain in .com, respectively), and its interaction with denial-of-existence proofs can reveal structural information of DNS zone. In particular, the transition from NODATA to NXDOMAIN responses suggest the presence of closest enclosers and enable inference of partial zone structure.
-
LongYiming Zhang (Tsinghua University), Hang Zhou (Tsinghua University), Tao Wan (CableLabs & Carleton University), Haixin Duan (Tsinghua University), Liqian Hao (China Unicom & Tsinghua University), Baojun Liu (Tsinghua University), Chaoyi Lu (Zhongguancun Laboratory), Jingying Zhou (China Unicom, Guangdong Branch)Abstract: Modern cellular networks rely on DNS-based service discovery to select internal control-plane and data-plane functions. Such queries are intended to remain within operators' private namespaces, yet misconfigurations can leak them to the public Internet. Despite warnings from 3GPP and IETF decades ago, the prevalence and security implications of such leakage remain unexamined. We present the first large-scale measurement of cellular service-discovery DNS leakage. By analyzing two days of B-root traffic for each year from 2021 to 2025 and identifying internal cellular queries via 3GPP naming patterns, we observed 13.46 million leaked queries from 139 countries, covering 135,845 unique FQDNs. Most leakage (93.10\%) was sporadic, suggesting operators remediated issues over time, yet a small set showed persistent leakage across all five years. Leakage frequently arose in cross-operator and cross-country scenarios, consistent with roaming behavior. While the increased load on B-root is minimal (0.02\% on average), leaked names often encode internal deployment parameters (e.g., base-station identifiers), exposing information that operators do not otherwise publish. Using open-source testbeds, we further show that an on-path adversary can redirect service-discovery responses once such queries leave the operator's network, which would otherwise be infeasible without the leak. Our work highlights the need for stricter DNS isolation in cellular networks.
-
LongYujia Zhu (Institute of Information Engineering, CAS), Linkang Zhang (Institute of Information Engineering, CAS), Baiyang Li (Institute of Information Engineering, CAS), Zhen Li (Institute of Information Engineering, CAS), Gang Xiong (Institute of Information Engineering, CAS), Qingyun Liu (Institute of Information Engineering, CAS), Xuebin Wang (Institute of Information Engineering, CAS)Abstract: The extension of DNS encryption toward the recursive-to-authoritative link (RFC 9539) introduces an opportunistic paradigm with an unexplored real-world security posture. Accurately measuring this ecosystem requires joint observation of resolver behaviors and authoritative capabilities. In this paper, we conduct a dual-perspective, Internet-scale measurement to characterize this emerging ecosystem. By synthesizing behaviors from both the authoritative supply side and the recursive client side, we surface critical operational realities and potential risks. Our empirical findings reveal three key insights: (i) encrypted authoritative deployment remains heavily concentrated among a small number of hosting providers, while the encrypted upstream behavior of resolvers is dominated by two popular public DNS providers; (ii) on both sides, deployed practice diverges from the design assumptions of the RFC regarding state management and privacy mechanisms; and (iii) these joint operational behaviors inherently leak residual privacy data despite nominal encryption, and expose authorities to slow-consumption connection exhaustion risks. Finally, we provide actionable recommendations for authoritative operators, resolver implementers, and the IETF to narrow the gap between standard intent and deployment reality.
-
ShortYevheniya Nosyk (KOR Labs), Simon Fernandez (Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG), Andrzej Duda (Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG), Maciej Korczyński (Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG)Abstract: DNS resolvers increasingly support various encryption protocols, ensuring their communication with end clients remains confidential to external observers. The recursive-to-authoritative link has long been overlooked though, despite multiple reports on traffic analysis and response injection by state censors. The experimental RFC 9539 addresses this confidentiality gap with a unilateral and opportunistic mechanism - recursive resolvers probe nameservers for DNS-over-TLS or DNS-over-QUIC support and, if successful, communicate over the encrypted channel. In this paper, we measure the deployment of RFC 9539 (ADoT/ADoQ) in the wild, covering both recursive resolvers and authoritative nameservers. We identify fewer than 1% (3.1 M) of registered domains supporting authoritative DoT or DoQ, with one provider accounting for the vast majority of these deployments. Ultimately, our data-driven study informs DNS operators that increasingly consider the deployment of authoritative DoT/DoQ but lack concrete numbers on the current state of deployment.
- 10:30 - 11:00 - Break
- 11:00 - 12:30 - Session 9
- Session 9: Transport Protocols and Real-Time Communication (Session Chair: TBA)
-
LongMarcel Kempf (Technical University of Munich), Nikolas Gauder (Technical University of Munich), Johannes Späth (Technical University of Munich), Georg Carle (Technical University of Munich), Johannes Zirngibl (Max Planck Institute for Informatics)Abstract: The Internet Engineering Task Force (IETF) specified QUIC in 2021, marking a significant shift in the Internet transport landscape. With its user-space implementation and flexible transport parameters, QUIC promised to enable rapid deployment of new extensions and features, fostering transport protocol innovation. While early studies captured the initial surge of deployment, long-term evolution and maturation of the QUIC ecosystem remain largely unexplored. In this paper, we present a comprehensive longitudinal study of QUIC based on a five-year dataset of weekly Internet-wide scans from April 2021 to February 2026. We use fingerprinting to identify 22 distinct libraries and track their market share over time. Further, we analyze the diversity and evolution of transport parameter configurations and the adoption of new extensions. Our analysis reveals a steady growth in QUIC adoption across both IPv4 and IPv6. Additionally, we show that while the ecosystem is changing, a significant “long tail” of outdated drafts and default configurations persists. This work provides the first multi-year overview of the QUIC landscape, offering insights into its volatility, the speed of protocol updates in user space, and the use of QUIC’s flexible transport parameters on the Internet.
-
LongZahra Yazdani (Georgia Tech), Fahad Hilal (Max Planck Institute for Informatics), Kevin Vermeulen (CNRS, Ecole Polytechnique), Cecilia Testart (Georgia Institute of Technology), Alberto Dainotti (Georgia Tech), Tiago Heinrich (Max Planck Institute for Informatics), Taha Albakour (Max Planck Institute for Informatics)Abstract: TCP Fast Open (TFO) is a TCP option designed to reduce connection establishment overhead. In this paper, we conduct the first large-scale investigation of TFO adoption on the Internet across both IPv4 and IPv6. Contrary to prior belief, we find a non-negligible rate of server-side adoption, particularly among large content provider networks. We observe similar enablement rates for IPv4 and IPv6, reaching 24% of our target addresses on port 443 (HTTPS). Further, we investigate top domains and find TFO to be more commonly enabled over IPv6 for dual-stack domains than in our broader Internet census. Our results reveal high TFO adoption among hypergiants, including Cloudflare (99%), Akamai (59%), and Google (40%). Additionally, we investigate middlebox-induced path impairments and find that interference with SYN packets affects few networks, suggesting that middleboxes pose only a limited obstacle to TFO operation. Lastly, we explore lightweight yet effective measurement opportunities that leverage server TFO cookies to measure the prevalence of aliased and anycast prefixes.
-
LongAnthony Gatti (University of Pittsburgh), Jiaao Lyu (University of Pittsburgh), Tanishka Deshpande (University of Pittsburgh), Amy Babay (University of Pittsburgh)Abstract: TCP slow-start's conservative ramp-up leaves short flows well below available bandwidth during their most performance-critical rounds. SUSS (SIGCOMM 2024) addresses this with a sender-side add-on to CUBIC that predicts when exponential growth can safely continue and accelerates the congestion window accordingly, reporting over 20\% flow-completion-time (FCT) improvement across 28 diverse scenarios. We replicate SUSS on an AWS-based testbed spanning six wired inter-region paths with RTTs from 8 to 262\,ms and one residential WiFi path. We further extend the original paper in two directions. First, we characterize the relationship between SUSS's benefit and RTT when the server region and last-hop link type are held fixed. We find a clear lower-bound effect below 50\,ms RTT --- the regime where the original paper's claim does not apply --- and substantial path-specific variation above it. Across our five wired scenarios where the claim applies, improvement averaged across flow sizes ranges from 27\% to 44\%, with no monotonic relationship to RTT. RTT alone is not a reliable predictor of SUSS's benefit. Second, we adapt SUSS's acceleration principle to BBR's STARTUP phase as an exploratory extension, finding that the principle transfers to a fundamentally different congestion control algorithm but at substantially smaller improvement magnitude (1--9\% versus 12--47\%). We believe this is consistent with BBR's already-aggressive baseline that leaves a narrower under-utilization window for acceleration to exploit. We additionally evaluate SUSS under concurrent long-lived background traffic and find that its FCT improvements persist on contended paths under both CUBIC and BBR.
-
LongSimon Sundberg (Karlstad University), Anna Brunstrom (Karlstad University), Simone Ferlin-Reiter (Red Hat and Karlstad University), Jesper Dangaard Brouer (Cloudflare), Toke Høiland-Jørgensen (Red Hat)Abstract: With networking moving into the sub-millisecond latency domain, latency in the end host itself can become a significant barrier to achieving consistently low application latency. Both the physical interconnect between the network card and the CPU, the kernel network stack, and the scheduling of applications themselves can be considerable sources of latency. Previous work has studied host latency at various levels, yet there remains a lack of methods and tools to continuously monitor host latency in production. To remedy this, we present netstacklat, a monitoring tool that captures latency at several points in the host network, from the early parts of the Linux kernel network stack all the way until the application reads the data. We evaluate netstacklat in a testbed, demonstrating its ability to capture host latency across 144 variations of HTTP workloads for Nginx and Apache, while also showing how the low monitoring overhead does not inflate tail latency by more than 6%, where previous monitoring solutions increase it by over 100%. Furthermore, we share our initial findings from deploying netstacklat in Cloudflare's global CDN network.
-
LongJonathan Binkle (Saarland University), Vaishnavi Raghavajosyula (Max Planck Institute for Informatics), Tiago Heinrich (Max Planck Institute for Informatics), Anja Feldmann (Max Planck Institute for Informatics), Tobias Fiebig (TU Wien), Johannes Zirngibl (Max Planck Institute for Informatics)Abstract: Efficient network transmission, especially in IPv6 since on-path fragmentation is not possible, depends on selecting the right packet size. The Maximum Transmission Unit (MTU) determines the largest packet size that can traverse a link and the Path MTU (PMTU) determines the size for an end-to-end path without fragmentation. Due to frequent failures in Path MTU Discovery (PMTUD), protocols must fall back to conservative packet sizes to ensure reachability, sacrificing potential throughput. This creates a need for inference-based approaches that estimate a suitable path size. Existing PMTU studies are outdated, limited in scope, and rely on indirect MSS inference or PMTUD-dependent signals, offering little insight into where MTU bottlenecks occur along the path. This paper presents a measurement study of PMTUs using RIPE Atlas to characterize their behavior in general, across protocol layers (IPv4/IPv6) and network locations thus covering 76,700 paths across access networks, the Internet core, and web-serving infrastructure. We observe 97 distinct PMTUs, though a small set:[1500 B (80.5%), 1492 B (11.7%), and 1480 B (1.5%)], dominates PMTU distributions differ across IP versions: while 99% of IPv4 paths support at least 1420 B, IPv6 paths commonly operate at 1280 B. We further show that sub-1500 B PMTUs are primarily located near network edges, while core paths rarely exhibit such constraints. Finally, we measure PMTU vs. Maximum Segment Size (MSS)-inferred PMTU mismatches in 18.5% of paths.
- 12:30 - 14:00 - Lunch
- 14:00 - 15:30 - Session 10
- Session 10: Blockchain and Decentralized Systems (Session Chair: TBA)
-
LongQiming Ye (Hong Kong University of Science and Technology (GZ)), Wen Yang (Hong Kong University of Science and Technology (GZ)), Saidu Sokoto (City St George’s, University of London), Leonhard Balduf (TU Darmstadt), Michał Król (City St George’s, University of London), Onur Ascigil (Lancaster University), Gareth Tyson (Hong Kong University of Science and Technology (GZ))Abstract: Decentralized social networks present independent alternatives to centralized counterparts, but there are open questions about their efficiency and long-term viability. Recently, a new class of hybrid social networks has emerged, which combines blockchain-based and off-chain operations to balance availability, performance, and sustainability. Farcaster is the first and largest hybrid social network that adopts this strategy. It combines on-chain identity management with peer-to-peer content dissemination via independent servers (aka hubs). In this paper, we present its first large-scale empirical study, analyzing 240 million user-generated messages (aka casts) across the platform and focusing on 3,000 independent hubs. We uncover asymmetries that challenge the platform's core goals of availability, performance, and sustainability. Many hubs exhibit repeated periods of non-responsiveness, and the network struggles to maintain a consistent state across hubs. Moreover, infrastructure centralization emerges due to hosting concentration and network topology skew. While gossip-based dissemination enables rapid initial delivery, achieving full consistency remains challenging: on average, 80% of hubs’ casts reach half of other hubs within 9.2 minutes, whereas the remaining 20% require up to 57 minutes. Finally, despite collecting over $2.6M in storage fees, there is no redistribution to hub operators, raising concerns about long-term sustainability. Our findings provide pointers for the design of hybrid and decentralized social platforms.
-
LongDennis Trautwein (University of Göttingen), Cornelius Ihle (University of Göttingen), Moritz Schubotz (FIZ Karlsruhe), Corinna Breitinger (University of Göttingen), Bela Gipp (University of Göttingen)Abstract: The promise of decentralized peer-to-peer (P2P) systems is fundamentally gated by the challenge of Network Address Translation (NAT) traversal, with existing solutions often reintroducing the very centralization they seek to avoid. This paper presents the first large-scale measurement study of a fully decentralized NAT traversal protocol, Direct Connection Upgrade through Relay (DCUtR), within the production libp2p-based InterPlanetary File System (IPFS) network. Drawing on over 4.4 million traversal attempts from 85,000+ distinct networks across 167 countries, we provide an empirical analysis of modern P2P connectivity. We establish a conditional success rate of 70% +- 7.1% for the hole-punching stage, given that prerequisite relay reservation and public address discovery succeed, providing a crucial new benchmark for the field. Critically, we empirically challenge the long-held belief of UDP's superiority for NAT traversal, demonstrating that DCUtR's high-precision, RTT-based synchronization yields statistically indistinguishable success rates for both TCP and QUIC (~70%). Our analysis further validates the protocol's design for permissionless environments by showing that success is independent of relay characteristics and that the mechanism is highly efficient, with 97.6% of successful connections established on the first attempt. Building on this analysis, we propose a concrete roadmap of protocol enhancements aimed at achieving universal connectivity and contribute our complete dataset to foster further research in this domain.
-
LongWenlong Zhang (Zhejiang University), Yajin Zhou (The Chinese University of Hong Kong), Lei Wu (Zhejiang University)Abstract: Cross-chain bridges are critical for asset mobility across fragmented blockchain ecosystems and the backbone of multi-chain DeFi. However, cross-chain transaction asynchrony introduces an underexplored risk: transaction stalling. Unlike active threats (e.g., security breaches), stalled transactions trap user funds, undermine liquidity, and distort transparency. Despite these consequences, they remain largely overlooked in existing research. This paper presents the first systematic study on cross-chain transaction stalling, addressing three key questions: how widespread stalling is, how it impacts cross-chain ecosystems, and why it occurs. We build a dataset covering nine mainstream cross-chain bridges through transaction matching with unique identifiers, cross-verification via bridge explorers and on-chain logs, and a lifecycle-based tracing framework to identify root causes. Our findings show notable asymmetries: OP Stack-based bridges have stalling ratios up to 41.33% for L2-to-L1 withdrawals (due to user-dependent workflows), while L1-to-L2 flows show near-zero stalling. Millions of dollars in assets are stranded across protocols, eroding liquidity and distorting proof-of-reserve (PoR) mechanisms, because stranded assets are not excluded from reserve calculations, which inflates reported ratios. For example, Base Bridge’s top 50 affected tokens see 0.2%–50% PoR distortion. Root causes include inadequate parameter validation, insufficient relayer incentives, and incomplete user actions. This work identifies cross-chain transaction stalling as a systemic risk, providing insights for improving cross-chain reliability.
-
LongPengfei Li (Zhejiang University), Lei Wu (Zhejiang University), Tianyang Chi (Beijing University of Posts and Telecommunications), Siwei Wu (City University of Hong Kong), Runhuai Li (BlockSec), Sophie Liu (Eigenphi), Cong Wang (City University of Hong Kong), Yajin Zhou (The Chinese University of Hong Kong)Abstract: Maximal Extractable Value (MEV) bots are integral to DeFi, yet their security is largely unexplored. From 203 reported attacks, we find that all incidents stem from access-control failures: about 80\% from missing initiator checks and the rest from improper checks. These patterns differ sharply from typical DeFi exploits and evade existing DApp-focused detection methods. To measure attacks at scale, we design a \textit{beneficiary-based} detection system and apply it to Ethereum data from 2022 to 2024. The system uncovers 698 previously unreported attacks against 215 MEV bots, causing over \$5.8M in losses. Losses and attacker revenues are highly skewed, with most damage concentrated among a few bots, and delayed bot responses enabling 483 secondary attacks. To evaluate latent risk, we build a \textit{simulation-based} exploitability analysis. We find that 9\% of MEV bots (638 out of 7,066) were profitably exploitable during 2022–2024, with potential losses up to \$45M. As of December 2024, 66 bots remain profitably exploitable, holding over \$0.2M in assets, and 23 continue to operate, posing ongoing risk.
-
LongMegha Sundriyal (Max Planck Institute for Security and Privacy), Jaehong Kim (KAIST), Wenchao Dong (Max Planck Institute for Security and Privacy), Wonjae Lee (Korea Advanced Institute of Science and Technology), Meeyoung Cha (Max Planck Institute for Security and Privacy (MPI-SP), Bochum, Germany)Abstract: The underground economy of the dark web is sustained not merely by markets but by resilient social groups that coordinate anonymously without institutional oversight. While prior research has focused on specific activities such as cryptomarkets, illicit services, and law enforcement interventions, the dark web as a broader social ecosystem remains largely underexamined. To bridge this gap, we present a large-scale measurement study of Dread, the most prominent dark web social forum, active since February 2018. We collect a large public-facing snapshot of Dread comprising 394,470 posts, 2,166,088 comments, 180,289 user profiles, and 1,999 communities, with content dating from February 2018 to January 2026. Our findings characterize Dread as a functionally specialized ecosystem, partitioned into discrete communities that serve specific interests, primarily centered on illicit trade and cybercrime. We also identify a sustained transition in financial discourse toward privacy-centric infrastructure, as users pivot from transparent ledgers (e.g., Bitcoin) to computationally intransigent protocols (e.g., Monero) in response to intensifying law enforcement. Underlying all of this is a structured, layered governance system, where content is regularly moderated, and PGP-encrypted messages are routinely embedded as private communication channels. These observations challenge the prevailing perception of dark web forums as chaotic and ungoverned. Instead, they reveal a highly organized, self-regulating ecosystem that rivals the organizational and administrative complexity of established surface web architectures.
- 15:30 - 16:00 - Break
- 16:00 - 17:30 - Session 11
- Session 11: BGP and Routing Security (Session Chair: TBA)
-
LongJustin Loye (IIJ Research Laboratory), Esteban Bautista (Univ. Littoral Cote d'Opale, LISIC), Romain Fontugne (IIJ Research Laboratory)Abstract: Internet disruptions are unpredictable events with potentially harmful societal and economic consequences. The Border Gateway Protocol (BGP) offers a valuable vantage point for their timely detection in the form of protocol-level anomalies. Yet, anomaly detection pipelines in BGP still suffer from severe limitations, such as needing large training datasets or relying on complex features that are hard to extract in real time. To address these issues, this paper introduces MADBGP, a novel anomaly detection pipeline that is based on (\emph{i}) a modeling of the Internet as a temporal graph; and (\emph{ii}) an anomaly detection algorithm that analyzes such temporal graph. Our contribution is threefold: (\emph{i}) we propose a new graph-based modeling of the AS-level Internet; (\emph{ii}) we develop and deploy QuickRIB: a scalable tool for generating temporal graphs from public BGP data; and (\emph{iii}) we extend MAD, a state-of-the-art method for anomaly detection in temporal graphs, to weighted temporal graphs and use it to detect real anomalies for the first time. In sum, MADBGP provides an unsupervised, multi-scale, and near real-time anomaly detection pipeline that is suitable for both researchers and network operators. Extensive experiments on ten major BGP events involving outages and route leaks at different scales demonstrate that MADBGP significantly outperforms IODA's BGP results.
-
LongThe Circuit Breaker That Cried Wolf: Measuring BGP Maximum-Prefix Exceedance and Connectivity ImpactOrlando E. Martínez-Durive (IMDEA Networks Institute), Antonis Chariton (Cisco ThousandEyes), Marco Fiore (IMDEA Networks Institute)Abstract: The BGP maximum-prefix limit serves as a network circuit breaker, tearing down sessions when announcements exceed a configured threshold. Despite its operational significance, little is known empirically about how accurately these limits reflect reality, how often they are exceeded, or what connectivity impact results from their exceedance. We present a first large-scale measurement study of BGP maximum-prefix limits in the wild, combining 12 months of RIPE RIS snapshots with PeeringDB metadata and CAIDA AS Rank data. Only 32.49% of RIS-visible ASes appear in PeeringDB, and 21--27% of recent limit fields carry a value of zero, indicating no declared limit; after filtering for recency and non-zero limits, 14,973 ASes form our analytical dataset. Among these, 14.55% (IPv4) and 11.24% (IPv6) exceed their declared limit intermittently, and 2.15% and 2.32% do so in every observed snapshot. These exceedances are driven by organic growth, such as new space acquisition or deaggregation, yet the mechanism responds identically to a route leak: Tier-1 and Major peers account for 18.4% of IPv4 lost sessions despite representing roughly 0.1% of all ASes. Based on our findings, we offer recommendations for operators and IXPs and call for renewed IETF standardization of dynamic limit negotiation.
-
LongWeitong Li (Virginia Tech), Yongzhe Xu (Virginia Tech), Mingwei Zhang (Cloudflare), Vasileios Giotsas (Cloudflare), Taejoong Chung (Virginia Tech)Abstract: The Resource Public Key Infrastructure (RPKI) and Route Origin Validation (ROV) are critical to securing Internet routing, yet measuring real-world ROV deployment and dependencies remains challenging. Existing methods (control-plane analysis and data-plane probing) often struggle to distinguish local ROV adoption from upstream filtering at scale and face diminishing visibility as ROV adoption grows. We present RScope, a measurement framework that dynamically manipulates Route Origin Authorizations (ROAs) from controlled publication points to isolate the impact of individual Relying Party (RP) servers. By selectively invalidating prefixes for specific RPs and probing connectivity changes, RScope maps dependencies between 21,827 ASes, their transit providers, and RP infrastructure. Our findings reveal that 1,127 ASes actively deploy ROV, while 1,815 rely on upstream protection, leaving 69.6% of such “protected” ASes vulnerable to local hijacks. Notably, large ISPs disproportionately influence global security—enabling ROV in the top 100 ASes protects 27.6% of networks, while disabling it in Tier-1s exposes 23.8% to hijacks. We also uncover operational delays: router enforcement lags ROA updates by 37 minutes on average. These results highlight systemic risks in incremental ROV adoption, emphasizing the need for resilient RP configurations and broader deployment. We open-source tools and datasets to support ongoing routing security efforts.
-
LongMauro Farina (University of Trieste), Martino Trevisan (University of Trieste), Alberto Bartoli (University of Trieste)Abstract: The Border Gateway Protocol (BGP) lacks native mechanisms to verify that routes adhere to authorized paths. This makes the global infrastructure vulnerable to route leaks, which have caused several major Internet disruptions. Autonomous System Provider Authorization (ASPA) is an emerging technology, currently at Internet-draft stage, that extends the Resource Public Key Infrastructure (RPKI) by allowing ASes to cryptographically declare their upstream providers; prior work has shown its effectiveness in detecting route leaks and certain forms of path manipulation. Yet, as ASPA moves toward standardization and Regional Internet Registries enable it in their RPKI dashboards, its actual deployment, operational dynamics, and alignment with independently inferred relationships remain unstudied. In this paper, we present the first measurement study of real-world ASPA deployment, tracing the evolution of 1646 adopting ASes from October 2023 to March 2026. We further introduce a methodology to apply ASPA AS Path verification---originally designed for live BGP routers---to passively collected BGP data, and use it to identify thousands of potential route leaks across nearly 57 million AS Paths. Finally, we compare ASPA-declared providers against two publicly available datasets of inferred providers, and find a recall up to 93%.
-
LongMalte Tashiro (IIJ Research Laboratory / SOKENDAI), Romain Fontugne (IIJ Research Laboratory), Kensuke Fukuda (NII / SOKENDAI), Randy Bush (IIJ Research Laboratory / Arrcus, Inc), Doug Madory (Kentik)Abstract: In 2023, Italy’s largest internet service provider was struck by a network outage which affected thousands of users and lasted nearly five hours. The root cause was a connectivity problem at its only upstream autonomous system (AS). This case highlights the fragility introduced by relying on a single upstream. In this paper we assess Internet resilience by quantifying single-upstream ASes and studying their main characteristics. First, we analyze BGP data and show that single-upstream ASes account for half of the ASes on the Internet. Comparison between IPv4 and IPv6 shows that dual stack ASes are more likely to rely on a single upstream only in IPv6. Furthermore, if they have dependencies in both IPv4 and IPv6, they usually apply the same networking policy. To validate these results, we perform active measurements to a random sample of 6366 ASes using traceroutes and BGP poisoning. We find that 72% of the ASes indeed rely on a single upstream, but also identify backup upstreams for 5%. Links to backup upstreams are normally not visible in BGP and only used if the primary upstream fails. We characterize the measured ASes using different size metrics and AS type classifications and find that backup upstreams are usually only employed by smaller ASes, whereas reliance on a single upstream is seen for both small and large ASes. Finally, we perform a case study of large eyeball networks covering 247 ASes from 163 countries for IPv4 and 141 ASes from 113 countries for IPv6. In total, 78% of analyzed countries in IPv4 and 81% in IPv6 have at least one large provider relying on a single upstream.
-
ShortPoornima Mani (Max-Planck Institute for Informatics), Q Misell (Max-Planck Institute for Informatics), Pascal Hennen (DE-CIX / Max-Planck Institute for Informatics), Anja Feldmann (Max-Planck Institute for Informatics), Tobias Fiebig (TU Wien)Abstract: The Border Gateway Protocol (BGP) enables the complex interconnections of the Internet, a network of networks --- each with its own perspective on the Global Routing Table (GRT). Route Collector (RC) projects like RouteViews and RIPE RIS collect these perspectives from networks and make that data publicly available; data broker services provide interfaces to these BGP archives. Researchers frequently use this data to study and understand the routing ecosystem. Although researchers expect these data sources to be flawed their reliability and consistency has yet to be studied in depth. We fill this gap by investigating the temporal (are routes recorded \emph{when} they should be) and internal (are routes recorded \emph{correctly}) consistency. Furthermore, we evaluate if BGPStream, a popular BGP RC data broker, reliably returns the data files it should. We find widespread inconsistencies for RC data, and share our approach to enable spot-checks when working with RC data. For BGPStream, we find regularly missing data since 2017, including \emph{all} files for RIPE RIS for the month of December 2022.
- 18:00 - 20:00 - Poster Session
- 9:00 - 10:30 - Session 12
- Session 12: Online Tracking and Privacy (Session Chair: TBA)
-
LongAbdullah Ghani (Lahore University of Management Sciences), Yash Vekaria (University of California, Davis), Zubair Shafiq (University of California, Davis)Abstract: Tracking pixels are used to optimize online ad campaigns through personalization, re-targeting, and conversion tracking. Past research has primarily focused on detecting the prevalence of tracking pixels on the web, with limited attention to how they are configured across websites. A tracking pixel may be configured differently on different websites. In this paper, we present a differential analysis framework: PixelConfig, to reverse-engineer the configurations of Meta Pixel deployments across the web. Using this framework, we investigate three types of Meta Pixel configurations: activity tracking (i.e., what a user is doing on a website), identity tracking (i.e., who a user is or who the device is associated with), and tracking restrictions (i.e., mechanisms to limit the sharing of potentially sensitive information). Using data from the Internet Archive’s Wayback Machine, we analyze and compare Meta Pixel configurations on 18K health-related websites with a control group of the top 10K websites from 2017 to 2024. We find that activity tracking features, such as automatic events that collect button clicks and page metadata, and identity tracking features, such as first-party cookies that are unaffected by third-party cookie blocking, reached adoption rates of up to 98.4%, largely driven by the Pixel’s default settings. We also find that the Pixel is being used to track potentially sensitive information, such as user interactions related to booking medical appointments and button clicks associated with specific medical conditions (e.g., erectile dysfunction) on health-related websites. Tracking restriction features, such as Core Setup, are configured on up to 34.3% of health websites and 8.7% of control websites. However, even when enabled, these tracking restriction features provide limited protection and can be circumvented in practice.
-
LongNicole Zagson (Northeastern University), Sarah Elizabeth Gillespie (Northeastern University), Jason Veara (Northeastern University), Devin Patel (Northeastern University), David Choffnes (Northeastern University), Alan Mislove (Northeastern University), Aanjhan Ranganathan (Northeastern University), Piotr Sapiezynski (Northeastern University), Christo Wilson (Northeastern University)Abstract: Modern vehicles typically integrate a wide array of sensors, software systems, and always-on connectivity, effectively becoming ``smartphones on wheels.'' Such vehicles are part of a broader ecosystem that includes companion mobile apps and third-party services, raising significant privacy risks due to the sensitivity of vehicular data. Despite documented harms, the community still lacks a good understanding of the privacy implications of the connected vehicle ecosystem. In this paper, we present the first large-scale empirical study of data privacy in the connected vehicle ecosystem. In collaboration with a consumer vehicle testing organization, we use two perspectives to better understand the data that is being collected and shared. First, we test 21 late model vehicles using a combination of Wi-Fi interception and controlled cellular capture under a range of conditions to observe the vehicles' network traffic patterns. Second, we instrument 30 vehicle companion apps paired with on-site vehicles to build a more complete picture of the data flows to third-party services. We find that both vehicles and companion apps contact numerous third-party domains, including advertisers and trackers. For example, all but two of the vehicles tested send traffic to at least one third party, and seven of 30 apps transmit sensitive identifiers to third-party companies, with wide variation across brands and models. Our findings underscore the need for continued measurement and scrutiny of the connected vehicle ecosystem.
-
LongXu Lin (Washington State University), Shujiang Wu (F5, Inc.), Jacob Anthony Eckfeldt (Washington State University), Pengfei Sun (F5, Inc.), Jason Polakis (University of Illinois Chicago)Abstract: In this paper, we present an in-depth investigation of a large dataset of browser fingerprints from users authenticating with high-value web services. Our analysis results in novel observations and insights that subvert our community’s established understanding of browser fingerprinting, by revealing the inherent instability of prevalent attributes even within stable execution environments. In other words, we identify patterns wherein fingerprints change even when no change has been made to the device (e.g., due to minor deviations in browsers’ floating point arithmetics or HTML element-size calculations). Naturally, this poses a major obstacle to the real-world effectiveness of browser-fingerprinting. To that end, we propose Predictive-FP, a novel adaptive fingerprinting strategy that leverages historical population-wide information to anticipate unstable fingerprinting values while also ignoring ephemeral attributes (e.g., version numbers). Our extensive experimental evaluation demonstrates how our system significantly outperforms the state of the art, especially over longer periods of time, and is able to re-identify∼51% more devices after 6 months. While our technique enables more robust authentication, it also presents a potential privacy risk. Accordingly, we develop a configurable privacy-preserving extension for users, that strategically modifies an extensive set of fingerprinting attributes to prevent tracking on non-allowlisted sites.
-
LongHyunSeok (Daniel) Jang (NORTHWESTERN UNIVERSITY), Hazem Ibrahim (New York University Abu Dhabi), Rohail Asim (New York University Abu Dhabi (NYUAD)), Matteo Varvello (Nokia), Yasir Zaki (New York University Abu Dhabi (NYUAD))Abstract: Bluetooth Low Energy (BLE) location trackers, or "tags", are popular consumer devices for monitoring personal items, which rely on their respective network of \textit{companion devices} to relay their location back to the owner. In 2023, Ibrahim et al. studied the performance of AirTag and SmartTag through controlled and in-the-wild experiments across six countries, finding that both tags achieve broadly similar performance, locating objects within 100 meters roughly 60\% of the time under 10 minutes. These experiments were conducted when COVID-19 social distancing measures were in practice, potentially reducing companion-device density and shaping the reported findings. In this work, we replicate the target paper with an improved data collection methodology and expand its scope by including Tile and scaling the in-the-wild campaign across 29 countries. Our controlled experiments reveal notable changes in tag behavior since the prior study, including a marked increase in AirTag's update rate. Our broader in-the-wild dataset highlights that companion device density is the primary determinant of tag performance, overshadowing technological differences between products. Compared to the results of the target paper, we observe higher combined accuracy across most reported conditions, with the largest gains at higher mobility speeds, which we discuss alongside the relaxation of pandemic-era social distancing.
-
ShortMichael Smith (Indiana University Bloomington), Riley Grossman (New Jersey Institute of Technology), Krzysztof Franaszek (Adalytics), Antonio Torres-Agüero (DeepSee.io), Pritam Sen (New Jersey Institute of Technology), Cristian Borcea (New Jersey Institute of Technology), Yi Chen (New Jersey Institute of Technology)Abstract: As third-party cookies fade because of browser restrictions, the online advertising ecosystem is turning to extended identifiers (EIDs) as an alternative. EIDs are persistent user identifiers, such as hashed email addresses, that are employed to link users across domains and devices. This paper presents a 41-month longitudinal study examining EID usage in over 145 million HTTP header bidding requests sent to six major supply-side platforms (SSPs) from 616,539 websites. Our findings show that EIDs are widely used and are becoming increasingly prevalent in the digital advertising ecosystem, reaching 83.76\% of studied websites by May 2025. Our analysis of the 18 popular EID providers that account for 99.42\% of all transmitted EIDs in our dataset raises concerns about the readiness of EIDs as an alternative to third-party cookie tracking. In terms of accuracy, only one identity provider consistently recognizes and identifies that the visitor is a self-identified bot crawler, and many providers regularly transmit multiple EIDs for the same visitor. We also identify privacy concerns with EIDs, as 12 of the providers create persistent EIDs that can identify the same user across visits, websites, devices, and months. Finally, we found that 16 providers transmit EIDs on EU websites without user consent.
-
LongVan Tran (University of Chicago), Shinan Liu (University of Hong Kong), Tian Li (University of Chicago), Nick Feamster (University of Chicago)Abstract: To address the scarcity and privacy concerns of network traffic data, various generative models have been developed to produce synthetic traffic. However, synthetic traffic is not inherently privacy-preserving, and the extent to which it leaks sensitive information, and how to measure such leakage, remain largely unexplored. This challenge is further compounded by the diversity of model architectures, which shape how traffic is represented and synthesized. We introduce a comprehensive set of privacy metrics for synthetic network traffic, combining standard approaches like membership inference attacks (MIA) and data extraction attacks with network-specific identifiers and attributes. Using these metrics, we systematically evaluate the vulnerability of different representative generative models and examine the factors that influence attack success. Our results reveal substantial variability in privacy risks across models and datasets. MIA success ranges from 0% to 88%, and up to 100% of network identifiers can be recovered from generated traffic, highlighting serious privacy vulnerabilities. We further identify key factors that significantly affect attack outcomes, including dataset diversity and how well the generative model fits the training data. These findings provide actionable guidance for designing and deploying generative models that minimize privacy leakage, establishing a foundation for safer synthetic network traffic generation. We release our code and evaluation framework to support reproducibility and facilitate future research.
- 10:30 - 11:00 - Break
- 11:00 - 12:30 - Session 13
- Session 13: Measurement Data Quality and Geolocation (Session Chair: TBA)
-
LongRobert Beverly (San Diego State University), Amreesh Phokeer (Internet Society), Oliver Gasser (IPinfo)Abstract: The five Regional Internet Registries (RIRs) provide the critical function of IP address resource delegation and registration. The accuracy of registration data directly impacts Internet operation, management, security, and optimization. In addition, the scarcity of IP addresses has brought into focus conflicts between RIR policy and IP registration ownership and use. The tension between a free-market based approach to address allocation versus policies to promote fairness and regional equity has resulted in court litigation that threatens the very existence of the RIR system. We develop WHEREIS, a measurement-based approach to geolocate delegated IPv4 and IPv6 prefixes at an RIR-region granularity and systematically study where addresses are used post-allocation and the extent to which registration information is accurate. We define a taxonomy of registration "geo-consistency" that compares a prefix's measured geolocation to the allocating RIR's coverage region as well as the registered organization's location. While in aggregate over 98% of the prefixes we examine are consistent with our geolocation inferences, there is substantial variation across RIRs and we focus on AFRINIC as a case study. IPv6 registrations are no more consistent than IPv4, suggesting that structural, rather than technical, issues play an important role in allocations. We solicit additional information on inconsistent prefixes from network operators, IP leasing providers, and collaborate with three RIRs to obtain validation. We further show that the inconsistencies we discover manifest in three commercial geolocation databases. By improving the transparency around post-allocation prefix use, we hope to improve applications that use IP registration data and inform ongoing discussions over in-region address use and policy.
-
ShortKatherine Izhikevich (UC San Diego), Ben Du (UCLA), Sumanth Rao (UC San Diego), Manda Tran (University of California, Los Angeles), Alisha Ukani (UC San Diego), Ray Bellis (Internet Systems Consortium), Liz Izhikevich (University of California, Los Angeles)Abstract: Geolocation plays a critical role in understanding the Internet. In this work, we analyze and fix operator-misreported geolocation. Using DNS root servers to detect speed of Internet violations, we conservatively infer that at least 2% of vantage points in the largest community-vantage point collection, RIPE Atlas, were not located at their operatorreported geolocation between January 2019 and April 2026. To increase the accuracy of future studies that use operatorreported geolocation, we open source our simple methodology, implement a reporting campaign of violating probes within RIPE, work with operators to fix the misreported geolocations, and release a continually updated dataset of RIPE vantage points that likely misreport geolocation.
-
LongSebastian Kappes (Max Planck Institute for Informatics), Anja Feldmann (Max Planck Institute for Informatics), Tobias Fiebig (TU Wien), Johannes Zirngibl (Max Planck Institute for Informatics)Abstract: Traceroute is an important Internet measurement tool used to infer the Internet’s physical, logical, and overlay topology as well as interfering devices, e.g., middlebox. It relies on network devices decreasing the Time to Live (TTL) in IP packet headers by one for each hop and sending Internet Control Message Protocol (ICMP) Time Exceeded error messages if the TTL reaches zero. However, we show that TTL jumps on the Internet exist, i.e., devices rewriting the TTL often to larger value, up to 255. Therefore, the remaining path after these rewrites is hidden from traceroutes resulting in incorrect inferences, e.g., inconsistently derived router and Autonomous System (AS) links. Based on controlled experiments and public data from RIPE Atlas and CAIDA Ark, we show that at least 47 ASes are impacted by at least one path impairing device rewriting the TTL. A prominent example is AT&T where more than 90 % of outgoing IPv6 paths from CAIDA Ark nodes are affected since 2023.
-
LongThomas Holterbach (bgproutes.io), Thomas Alfroy (bgproutes.io), Abbas Mohsenpour (UCLouvain), Thomas Krenc (IIJ Research Laboratory), kc Claffy (CAIDA), Cristel Pelsser (UCLouvain)Abstract: BGP data collected from operational routers by platforms such as RIPE RIS or RouteViews is essential to address connectivity issues and analyze the global Internet’s structure. However, the vast and ever-growing volume of this data makes deriving useful insights challenging: Existing APIs and dashboards offer limited perspectives, while analyzing large archives of MRT files is slow and cumbersome. We present ChatBGP, a domain-specific chatbot that turns plain-English questions about BGP data into Python code that runs swiftly and returns accurate answers. ChatBGP is enabled by four contributions: (i) a measurement study showing that BGP data is highly redundant and therefore highly compressible; (ii) a database scheme that exploits this redundancy to compress BGP data while enabling fast retrieval of specific data elements; (iii) an expressive API that exposes this database through a simple interface; and (iv) a prompt-engineering scheme that guides ChatGPT to synthesize optimized Python code that uses this API to accurately answer user queries, while retrieving only the data needed to answer each query. With ChatBGP, gaining insights into BGP data becomes effortless and significantly faster—e.g., delivering results in seconds instead of hours with existing tools. Network operators can rapidly diagnose connectivity issues, including security threats, researchers can analyze larger datasets while improving reproducibility, and students can engage with BGP monitoring in a more intuitive and interactive way.
-
LongTimofey Shpakov (ETH Zürich), Martin Burkhart (armasuisse Science and Technology), Laurent Vanbever (ETH Zürich)Abstract: Several recent studies claim that encrypted network traffic exhibits statistically significant randomness deficiencies due to differences in cryptographic algorithms, implementations, and parameterization, and that such patterns can be exploited for application classification in TLS traffic. This interpretation has been adopted by measurement researchers, most prominently in the often-cited ET-BERT study. We present a systematic replication of these randomness experiments, re-implementing them using the NIST Statistical Test Suite on encrypted traffic generated under controlled conditions. In the replicated work, the suite is not itself a detection component, but is used to argue that detectable patterns exist in ciphertext and thereby motivate the method. It is this evidentiary step that we examine. Our evaluation shows that the reported statistical signals do \textit{not} provide sufficient evidence of non-randomness beyond what is expected from statistical hypothesis testing and finite-sample effects. The observed test failures and discriminative patterns are consistent with methodological artifacts, including improper aggregation, p-value misinterpretation, and information leakage between training and test data. Overall, our findings suggest that assumptions about imperfect randomness in encrypted network traffic are unsubstantiated. Our critique is limited to randomness-based classification, and we make no claim about methods exploiting metadata, packet sizes, timing, or other side-channel information.
- 12:30 - 14:00 - Lunch
- 14:00 - 15:30 - Session 14
- Session 14: Access Networks: Satellite, Cellular, and Broadband (Session Chair: TBA)
-
Abstract: Measuring Starlink aviation performance at scale is challenging, as traditional methods require ``inside-out'' measurements conducted onboard aircraft. We present a scalable measurement methodology that characterizes the latency performance of Starlink-equipped commercial aircraft from the ground, without onboard access. Our approach leverages Starlink's structured IPv6 addressing scheme and the unique ``anycast users'' behavior. Each onboard Starlink router is assigned a globally reachable IPv6 address, and on anycast-capable aircraft, it remains stable across Point-of-Presence (PoP) boundaries, while associated MPLS labels are updated to reflect the current PoP association. We validate our ``outside-in'' measurements against inside-out measurements conducted during actual flights, confirming the fidelity of the ground-based approach. Applying our methodology across multiple airlines, we discovered more than 2,200 IPv6 anycast router addresses, identified over 900 aircraft from 16 airlines, and collected over 65 million \texttt{traceroute} records over five-month. Continuous latency measurements from 27 vantage points provide a global view of Starlink aviation performance, including latency, PoP handover behavior, and the impact of geographic factors, enabling the first worldwide characterization of Starlink in-flight connectivity (IFC) performance. Our methodology complements traditional inside-out measurements by reducing the logistical complexity of repeated data collection and enabling long-term measurements of Starlink aviation performance across the globe.
-
ShortJohan Garcia (Karlstad University), Simon Sundberg (Karlstad University), Anna Brunstrom (Karlstad University and University of Malaga)Abstract: In all networking systems, queuing is important to ensure appropriate resource utilization in the presence of bursty traffic and varying traffic demands. The Starlink access network is additionally also dynamic in terms of the capacity it can provide, and thus queuing plays an even greater role to ensure appropriate communication performance for the end-users while maintaining high resource utilization. However, for Starlink most system design details, along with the setup of the internal queuing, is private information and not publicly available. To address this we have developed a high-precision, burst-pattern controlled, traffic generation approach allowing us to precisely measure the one-way delay for Starlink. By analyzing the delay and loss in conjunction with a queue simulator we find that Starlink does not employ per-flow fair queuing or drop-tail buffers, but it does use drop-front buffer management. While drop-front reduces delay, it may also interfere with the assumptions made by loss-based congestion controls, potentially contributing to throughput degradation.
-
ShortJustus Fries (Technical University of Munich), Yannis Matezki (Technical University of Munich), Rohan Bose (Technical University of Munich), Nitinder Mohan (TU Delft)Abstract: SpaceX's Starlink has emerged as one of the largest network operators worldwide, serving over ten million subscribers across more than 150 countries. Real-time applications such as video conferencing, cloud gaming, and teleoperated control increasingly run over it. Yet how WebRTC's congestion controllers respond to Starlink's sub-second dynamics, namely its 15-second reconfigurations and satellite handovers, remains poorly understood. We present the first cross-layer measurement of WebRTC over Starlink, instrumenting Google Congestion Control (GCC) in a custom LibWebRTC testbed and pairing it with browser-side measurements of Microsoft Teams. Our campaign covers three vantage points in Europe and one in Antarctica at 2.5 and 10 Mbps targets. Even at Teams' conservative 2.5 Mbps uplink cap, the target is missed up to 34.9% of the time. Using browser metrics alone, we attribute up to 46% of these below-target periods to Starlink reconfigurations. GCC's delay-based estimator, which controls the bandwidth estimate 96–99% of the time, is the dominant pathway by which Starlink reaches the applications. Its overuse transitions cluster tightly around reconfigurations, while loss-based transitions remain uniformly distributed even at 10 Mbps. Satellite handovers, distinct from reconfigurations, cause longer-lasting overuse and the most aggressive sending-rate reductions in our dataset. Together, these findings reveal that GCC's conservatism, not intrinsic loss on the link, is the primary cost of Starlink's structural perturbations to real-time video.
-
LongZeyu Li (Northeastern University), Yufei Feng (Northeastern University), Dimitrios Koutsonikolas (Northeastern University)Abstract: Mobile virtual network operators (MVNOs) provide service over host mobile network operators (MNOs), yet users often assume that host carriers receive better performance. This paper presents, to our knowledge, the first simultaneous measurement study of MVNO–host MNO throughput competition with radio-side visibility. Across three operator families, five competition scenarios, and three U.S. cities, we run simultaneous TCP throughput tests over commercial subscriptions and collect cross-layer indicators including QCI, PCI, RAT mode, SA/NSA mode, resource block allocation, MCS, and MAC layer NR/LTE throughput. Our results show that host advantage is conditional rather than universal, and that no single factor consistently explains host–MVNO throughput ordering. Across operator families, observed performance differences are associated with different combinations of QoS class, RAT attachment, serving-cell context, radio-resource allocation, frequency band, and SA/NSA architecture. Under identical radio contexts, QoS-class ordering often coincides with down-link performance ordering. In contrast, when the radio contexts are different, RAT mismatch, cell attachment, or NSA LTE aggregation dominates the observed outcome. City-level comparisons further show that operator ordering can change across locations as RAT attachment, serving-cell selection, bandwidth, and radio conditions vary. Thus, host–MVNO performance must be interpreted through the joint radio and access context rather than from operator identity or any single QoS indicator.
-
LongJonatas Marques (IPinfo), Jared N. Schachner (University of Southern California), Nicole P. Marwell (University of Chicago), Nick Feamster (University of Chicago)Abstract: Crowdsourced datasets are vital for analyzing variation in Internet access performance across geographic areas and addressing these spatial disparities. Prior research using these data and pursuing similar objectives often examines performance variation between single spatial units, such as census tracts, and scrutinizes a limited set of sociodemographic variables like race or class composition to explain it. However, this approach may be insufficiently precise to characterize the patterns of, and explanations for, spatial differences in Internet performance. We argue that multilevel, multivariate models better represent Internet performance as a spatial phenomenon; these models decompose its variance at multiple spatial scales and permit inclusion of multiple explanatory factors at each level. To demonstrate the utility of this approach, we use multilevel, multivariate models to analyze spatial patterns of Internet latency (idle and under load) and jitter drawn from crowdsourced Ookla Speedtest data collected between 2022 and 2023. Despite prior research's emphasis on neighborhood variation in Internet performance, our multilevel models on crowdsourced data reveal that latency and jitter varies far more within neighborhoods than between them. Moreover, demographic differences in residential populations explain only a small portion of the modest neighborhood-level variance in Internet performance we estimate. Variation in infrastructure across counties and states appears to stratify performance to a far greater extent.
-
LongShuyue Yu (Columbia University), Bruce Spang (Netflix), Carson Garland (Columbia University), Kevin Chen (YouTube / Netflix), Ramki Gummadi (Google DeepMind), Gil Zussman (Columbia University), Ethan Katz-Bassett (Columbia University / Google), Te-Yuan Huang (Netflix), Renata Teixeira (Netflix)Abstract: Residential access networks today often appear well provisioned on paper, with high access speeds, modern Wi‑Fi, and content served from nearby caches. We investigate how such networks actually behave using packet‑level traces from 33 residential buildings owned and managed by Columbia University, hosting over 1,500 units. Our methodology infers TCP loss and access RTTs to detect congestion and estimates per‑packet in‑flight intervals to identify when packets from different services overlap on shared buffers. Even with low average utilization, we find that short traffic bursts can cause substantial loss and queueing delay, and that services often share capacity primarily with their own traffic. Pacing mitigates this burst-induced congestion by adding delay between packets and has been studied by YouTube and Netflix separately, but the vantage point of a single service cannot observe the effects of pacing on other services. We evaluate pacing through a real‑world experiment in which YouTube and Netflix alternately enable and disable pacing for users in this network. Under pacing, each service roughly halves its own median loss rate and reduces median access queueing delay by about 30–40%, while sustaining similar traffic volumes. For other services whose packets directly overlap paced video, median loss and queueing delay also decrease, even though pacing slightly increases the probability of such overlaps. Our results provide the first in-the-wild evidence that pacing is friendly in such a well-provisioned network.
- 15:30 - 15:45 - Concluding Remarks