<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>papers &#8211; Penn Database Group</title>
	<atom:link href="https://db.cis.upenn.edu/category/papers/feed/" rel="self" type="application/rss+xml" />
	<link>https://db.cis.upenn.edu</link>
	<description>Inventing the future of data management!</description>
	<lastBuildDate>Thu, 22 Jan 2026 18:39:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://db.cis.upenn.edu/wp-content/uploads/2022/02/cropped-simplified-shield-final-5-1-32x32.png</url>
	<title>papers &#8211; Penn Database Group</title>
	<link>https://db.cis.upenn.edu</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Penn DB Group Wins CIDR &#8217;26 Best Paper Award</title>
		<link>https://db.cis.upenn.edu/2026/01/22/penn-db-group-wins-cidr-26-best-paper-award/</link>
		
		<dc:creator><![CDATA[Ryan Marcus]]></dc:creator>
		<pubDate>Thu, 22 Jan 2026 18:38:11 +0000</pubDate>
				<category><![CDATA[awards]]></category>
		<category><![CDATA[papers]]></category>
		<guid isPermaLink="false">https://db.cis.upenn.edu/?p=655</guid>

					<description><![CDATA[Our group won the 2026 CIDR Best Paper Award for our paper, &#8220;Survivorship Bias in Industrial Database Workloads.&#8221; As we look at our workload logs now, how can we possibly predict the query the user wants to run, but cannot run? Marcus et al., &#8220;Survivorship Bias in Industrial Database Workloads&#8221;<a class="moretag" href="https://db.cis.upenn.edu/2026/01/22/penn-db-group-wins-cidr-26-best-paper-award/"> Read more</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">Our group won the 2026 CIDR Best Paper Award for our paper, &#8220;<a href="https://rm.cab/survivorshipbias">Survivorship Bias in Industrial Database Workloads</a>.&#8221; </p>



<figure class="wp-block-pullquote"><blockquote><p>As we look at our workload logs now, how can we possibly predict the query the user wants to run,  but cannot run?</p><cite>Marcus et al., &#8220;Survivorship Bias in Industrial Database Workloads&#8221;</cite></blockquote></figure>



<figure class="wp-block-image size-large"><img fetchpriority="high" decoding="async" width="1024" height="768" src="https://db.cis.upenn.edu/wp-content/uploads/2026/01/DSCF00422-1024x768.jpg" alt="Ryan Marcus receiving the best paper award from Nesime Tatbul and Sam Madden." class="wp-image-656" srcset="https://db.cis.upenn.edu/wp-content/uploads/2026/01/DSCF00422-1024x768.jpg 1024w, https://db.cis.upenn.edu/wp-content/uploads/2026/01/DSCF00422-300x225.jpg 300w, https://db.cis.upenn.edu/wp-content/uploads/2026/01/DSCF00422-768x576.jpg 768w, https://db.cis.upenn.edu/wp-content/uploads/2026/01/DSCF00422-1536x1152.jpg 1536w, https://db.cis.upenn.edu/wp-content/uploads/2026/01/DSCF00422-2048x1535.jpg 2048w" sizes="(max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Ryan received the award from Nesime Tatbul and Sam Madden at the conference.</figcaption></figure>



<div class="wp-block-media-text is-stacked-on-mobile"><figure class="wp-block-media-text__media"><img decoding="async" width="916" height="1108" src="https://db.cis.upenn.edu/wp-content/uploads/2026/01/paper.png" alt="" class="wp-image-657 size-full" srcset="https://db.cis.upenn.edu/wp-content/uploads/2026/01/paper.png 916w, https://db.cis.upenn.edu/wp-content/uploads/2026/01/paper-248x300.png 248w, https://db.cis.upenn.edu/wp-content/uploads/2026/01/paper-847x1024.png 847w, https://db.cis.upenn.edu/wp-content/uploads/2026/01/paper-768x929.png 768w" sizes="(max-width: 916px) 100vw, 916px" /></figure><div class="wp-block-media-text__content">
<p class="wp-block-paragraph">The <a href="https://rm.cab/survivorshipbias">paper</a> argues that workload traces observed in industrial settings represent a negotiation between the data platform and the platform&#8217;s users. Users mold their queries to run well on the platform, and engineers tune the platform to better meet user demands. This cycle, while great for both users and data platforms, creates a <em>survivorship bias </em>in the observed workloads: the most frequently processed queries are precisely the queries that are already working well.</p>
</div></div>



<p class="wp-block-paragraph">The paper&#8217;s authors include <a href="https://rmarcus.info">Ryan Marcus</a> (assistant professor), <a href="https://www.speculative.tech/">Jeffrey Tao</a> (PhD student), <a href="https://www.cis.upenn.edu/~pagewu/">Peizhi Wu</a> (PhD student alumnus), and <a href="https://zijie.me/">Zijie Zhao</a> (PhD student).</p>



<p class="wp-block-paragraph"></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Penn at VLDB 2025</title>
		<link>https://db.cis.upenn.edu/2025/08/22/penn-at-vldb-2025/</link>
		
		<dc:creator><![CDATA[Zack Ives]]></dc:creator>
		<pubDate>Fri, 22 Aug 2025 20:32:17 +0000</pubDate>
				<category><![CDATA[events]]></category>
		<category><![CDATA[papers]]></category>
		<guid isPermaLink="false">https://db.cis.upenn.edu/?p=615</guid>

					<description><![CDATA[This year, at VLDB 2025, Penn will be well-represented with a variety of papers. CausalMesh: A Causal Cache for Stateful Serverless Computing: Haoran Zhang (University of Pennsylvania); Shuai Mu (Stony Brook University); Sebastian Angel (University of Pennsylvania); Vincent Liu (University of Pennsylvania). In stateful serverless computing, workflows are broken into<a class="moretag" href="https://db.cis.upenn.edu/2025/08/22/penn-at-vldb-2025/"> Read more</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">This year, at VLDB 2025, Penn will be well-represented with a variety of papers.</p>



<p class="wp-block-paragraph"><strong>CausalMesh: A Causal Cache for Stateful Serverless Computing</strong>: Haoran Zhang (University of Pennsylvania); Shuai Mu (Stony Brook University); Sebastian Angel (University of Pennsylvania); Vincent Liu (University of Pennsylvania).  <em>In stateful serverless computing, workflows are broken into functions that may run on different physical machines, each with its own local cache. This distribution can lead to consistency errors, where one function reads stale data from its cache because a previous function in the same workflow wrote an update to a different machine&#8217;s cache. To solve this, researchers at Penn and Stony Brook developed <strong>CausalMesh</strong>, a novel caching system that guarantees &#8220;causal consistency,&#8221; ensuring operations are seen in a logical, cause-and-effect order across all machines. A key innovation of CausalMesh is that it provides this guarantee for most read/write operations without requiring costly coordination between servers or aborting transactions. As a result, CausalMesh delivers lower latency and higher throughput than existing approaches, enabling faster and more reliable state management in serverless applications.</em></p>



<p class="wp-block-paragraph"><strong><a href="https://vldb.org/pvldb/volumes/18/paper/A%20Practical%20Theory%20of%20Generalization%20in%20Selectivity%20Learning">A Practical Theory of Generalization in Selectivity Learning</a></strong>: Peizhi Wu (University of Pennsylvania), Haoshu Xu (University of Pennsylvania), Ryan Marcus (University of Pennsylvania), Zack Ives (University of Pennsylvania).  <em>This research provides a theoretical understanding of machine learning models used for query optimization in databases. While these models perform well in practice, there has been a significant gap in explaining <em>why</em> they work, especially when they encounter new or different queries (&#8220;out-of-distribution&#8221; or OOD) than those they were trained on. The paper bridges this gap by establishing the first theoretical guarantees for how these models generalize to OOD queries. Based on these new insights, the authors developed practical strategies that significantly improve the accuracy and real-world performance of existing models on unseen query types, making them more robust and reliable without sacrificing their original performance.</em></p>



<p class="wp-block-paragraph"><strong><a href="https://vldb.org/pvldb/volumes/18/paper/Holistic%20query%20Approximation%20via%20RL%20Modeling">Holistic query Approximation via RL Modeling</a></strong>. Susan Davidson (University of Pennsylvania), Tova Milo (Tel Aviv University), Kathy Razmadze (Tel Aviv University), Gal Zeevi (Tel Aviv University). <em>To accelerate slow queries during data exploration on large databases, researchers at Tel Aviv University and Penn have developed <strong>HARLM</strong>, a novel system for approximate query processing. While existing methods speed up aggregate queries (like <code>COUNT</code> or <code>AVG</code>) by using data samples, they fail to support non-aggregate queries that retrieve specific rows. HARLM presents a holistic solution by using Reinforcement Learning to identify an optimized, smaller subset of the data that works for both query types. This approach effectively learns to create a representative data sample that maximizes query accuracy while dramatically reducing execution time. Experiments show that HARLM significantly outperforms baseline methods, improving result accuracy by 30% and providing a 10-35x speedup.</em></p>



<p class="wp-block-paragraph"><strong>SHARQ: Explainability Framework for Association Rules on Relational Data</strong>: Hadar Ben‑Efraim (Bar-Ilan University), Susan B. Davidson (University of Pennsylvania), Amit Somech (Bar-Ilan University). <em>Association rule mining is a widely used technique for discovering patterns (e.g., &#8220;customers who buy X also buy Y&#8221;) in large datasets. However, a major challenge has been to quantify the actual importance of an individual data element, like &#8220;X,&#8221; to the entire set of rules it participates in. This paper introduces <strong>SHARQ</strong>, a novel method that uses Shapley values, a concept from cooperative game theory, to fairly and accurately measure the contribution of each element. While a naive calculation would be exponentially slow, the researchers developed highly efficient algorithms that compute this score in near-linear time. This breakthrough makes it practical to rank data elements, entire rules, and even attributes by their influence, providing a powerful new tool for explaining and gaining deeper insights from mined data.</em></p>



<p class="wp-block-paragraph"><strong><a href="https://vldb.org/pvldb/volumes/18/paper/Data-Agnostic%20Cardinality%20Learning%20from%20Imperfect%20Workloads">Data-Agnostic Cardinality Learning from Imperfect Workloads</a></strong>: Peizhi Wu (University of Pennsylvania), Rong Kang (ByteDance);Tieying Zhang (Bytedance), Jianjun Chen (Bytedance), Ryan Marcus (University of Pennsylvania), Zack Ives (University of Pennsylvania). <em>The authors, at Bytedance and Penn, have developed a new system called <strong>GRASP</strong> for cardinality estimation, a crucial task in database query optimization. Traditional methods need direct access to data, which is often restricted, while existing learning-based approaches struggle with the incomplete and imbalanced query workloads found in real-world scenarios. GRASP is a <strong>data-agnostic</strong> system specifically designed for these imperfect conditions. It uses a novel compositional design that allows it to generalize to new queries and is robust to skewed training data. By effectively modeling data distributions and join correlations without seeing the underlying data, GRASP consistently outperforms other query-driven models and, remarkably, can even match or exceed the accuracy of traditional methods that have full data access.</em></p>



<p class="wp-block-paragraph">(AIDB Workshop) <a href="https://api.zotero.org/users/3604318/publications/items/KNCCRRRJ/file/view"><strong>Exploring Wavelet Trees as Space-Efficient Physical-to-Sorted Mapping for Learned Indexes</strong>.</a> Anwesha Saha (Boston University), Aneesh Raman (Boston University), Ryan Marcus (University of Pennsylvania), Manos Athanassoulis (Boston University). <em>This paper explores Wavelet Trees as a compact way to map data between its physical and sorted order for learned indexes, which use machine learning models to replace traditional B+-tree nodes. The authors introduce Integer Wavelet Trees (IWTs), which significantly reduce memory usage—up to 84% less than B+-trees—but initially suffer from slow lookups due to cache inefficiencies. To address this, they propose T-way IWTs, which improve lookup speed while maintaining space efficiency, achieving 46% smaller memory footprints and 12% faster lookups compared to B+-trees. This study lays the groundwork for future designs, including their new idea of constellation maps, aimed at balancing speed and memory for learned index mappings.</em> This paper was a best paper honorable mention at the workshop!</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Penn at SIGMOD 2025</title>
		<link>https://db.cis.upenn.edu/2025/06/16/penn-at-sigmod-2025/</link>
		
		<dc:creator><![CDATA[Zack Ives]]></dc:creator>
		<pubDate>Mon, 16 Jun 2025 12:36:01 +0000</pubDate>
				<category><![CDATA[events]]></category>
		<category><![CDATA[papers]]></category>
		<guid isPermaLink="false">https://db.cis.upenn.edu/?p=582</guid>

					<description><![CDATA[The Penn Database and Data Systems Group is well-represented at SIGMOD 2025! At the aiDM workshop, co-chaired by our own Ryan Marcus, there are two papers: At the main SIGMOD conference, the following papers will be presented. Low Rank Learning for Offline Query OptimizationZixuan Yi (University of Pennsylvania)*; Yao Tian<a class="moretag" href="https://db.cis.upenn.edu/2025/06/16/penn-at-sigmod-2025/"> Read more</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The Penn Database and Data Systems Group is well-represented at SIGMOD 2025!</p>



<p class="wp-block-paragraph">At the <a href="http://www.aidm-conf.org/">aiDM workshop</a>, co-chaired by our own Ryan Marcus, there are two papers: </p>



<ul class="wp-block-list">
<li><strong><em>SERAG: Self-Evolving RAG System for Query Optimization,</em></strong>&nbsp;Hanwen Liu, Qihan Zhang, University of Southern California, Ryan Marcus, University of Pennsylvania, Ibrahim Sabek, University of Southern California.</li>



<li> <strong><em>Data-driven Adaptive Processing of Streaming ML Queries</em></strong>, by Phillip Hilliard, Rajeev Alur, Zachary Ives, University of Pennsylvania. This paper describes an adaptive query processing technique targeted at stream systems that incorporate machine learning components. When given a set of alternative machine learning models with different cost-accuracy trade-offs, it dynamically chooses the model that maximizes accuracy while satisfying a budgetary or quality-of-service constraint.</li>
</ul>



<p class="wp-block-paragraph">At the main SIGMOD conference, the following papers will be presented.</p>



<div class="wp-block-media-text is-stacked-on-mobile" style="grid-template-columns:34% auto"><figure class="wp-block-media-text__media"><img decoding="async" width="1024" height="888" src="https://db.cis.upenn.edu/wp-content/uploads/2025/06/limeqo_border_small-1024x888.png" alt="" class="wp-image-596 size-full" srcset="https://db.cis.upenn.edu/wp-content/uploads/2025/06/limeqo_border_small-1024x888.png 1024w, https://db.cis.upenn.edu/wp-content/uploads/2025/06/limeqo_border_small-300x260.png 300w, https://db.cis.upenn.edu/wp-content/uploads/2025/06/limeqo_border_small-768x666.png 768w, https://db.cis.upenn.edu/wp-content/uploads/2025/06/limeqo_border_small.png 1055w" sizes="(max-width: 1024px) 100vw, 1024px" /></figure><div class="wp-block-media-text__content">
<p class="wp-block-paragraph"><strong><a href="https://rm.cab/limeqo">Low Rank Learning for Offline Query Optimization</a></strong><br>Zixuan Yi (University of Pennsylvania)*; Yao Tian (The Hong Kong University of Science and Technology); Zack Ives (University of Pennsylvania); Ryan Marcus (University of Pennsylvania). </p>



<p class="wp-block-paragraph">This paper develops a novel technique based on low-rank matrix factorization, which allows a query optimizer to predict which query processing strategies will be useful for one query, based on performance of other queries.</p>
</div></div>



<div class="wp-block-media-text is-stacked-on-mobile" style="grid-template-columns:33% auto"><figure class="wp-block-media-text__media"><img loading="lazy" decoding="async" width="1024" height="888" src="https://db.cis.upenn.edu/wp-content/uploads/2025/06/bayesqo_border_small-1024x888.png" alt="" class="wp-image-597 size-full" srcset="https://db.cis.upenn.edu/wp-content/uploads/2025/06/bayesqo_border_small-1024x888.png 1024w, https://db.cis.upenn.edu/wp-content/uploads/2025/06/bayesqo_border_small-300x260.png 300w, https://db.cis.upenn.edu/wp-content/uploads/2025/06/bayesqo_border_small-768x666.png 768w, https://db.cis.upenn.edu/wp-content/uploads/2025/06/bayesqo_border_small.png 1055w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure><div class="wp-block-media-text__content">
<p class="wp-block-paragraph"><strong><a href="https://rm.cab/bayesqo">Learned Offline Query Planning via Bayesian Optimization</a></strong><br>Jeffrey Tao; Natalie Maus; Haydn Jones; Yimeng Zeng; Jacob Gardner; Ryan Marcus.</p>



<p class="wp-block-paragraph">Targeting queries that are going to be executed thousands of times, we propose an offline query optimizer that searches a wide variety of plans and incorporates query execution as a primitive. Our offline query optimizer combines variational auto-encoders with Bayesian optimization to find optimized plans for a given query.</p>
</div></div>



<p class="wp-block-paragraph"></p>



<ul class="wp-block-list">
<li><strong>SHARQ: Explainability Framework for Association Rules on Relational Data</strong><br>Hadar Ben Efraim (Bar-Ilan University); Susan B Davidson (University of Pennsylvania); Amit Somech (Bar-Ilan University)*. Association rules are an important technique for gaining insights over large relational datasets. However, it is difficult to explain the relative importance of data elements with respect to the rules in which they appear. This paper develops a measure of an element&#8217;s contribution to a set of association rules based on Shapley values, denoted SHARQ (ShApley Rules Quantification).</li>



<li><strong>Physical Visualization Design: Decoupling Interface and System Design</strong><br>Yiru Chen (Columbia University)*; Xupeng Li (Columbia University); Jeffrey Tao (University of Pennsylvania); Alana Ramjit (Cornell Tech); Ravi Netravali (Princeton University); Subrata Mitra (Adobe Research); Aditya Parameswaran (University of California, Berkeley); Javad Ghaderi (Columbia University); Dan Rubenstein (Columbia University); Eugene Wu (Columbia University)</li>



<li><strong>CARINA: An Efficient CXL-Oriented Embedding Serving System for Recommendation Models</strong><br>Peiqi Yin (The Chinese University of Hong Kong)*; Qihui Zhou (CUHK); Xiao Yan (Centre for Perceptual and Interactive Intelligence (CPII) ); Chao Wang (The Chinese University of Hong Kong); Eric Lo (Chinese University of Hong Kong); Changji Li (CUHK); Lan Lu (University of Pennsylvania ); Hua Fan (Alibaba Cloud); Wenchao Zhou (Alibaba Group); Ming-Chang YANG (The Chinese University of Hong Kong); James Cheng (CUHK)</li>
</ul>



<p class="wp-block-paragraph">At the demo sessions:</p>



<div class="wp-block-media-text is-stacked-on-mobile"><figure class="wp-block-media-text__media"><img loading="lazy" decoding="async" width="1024" height="771" src="https://db.cis.upenn.edu/wp-content/uploads/2025/06/penn_demo-1024x771.jpg" alt="" class="wp-image-608 size-full" srcset="https://db.cis.upenn.edu/wp-content/uploads/2025/06/penn_demo-1024x771.jpg 1024w, https://db.cis.upenn.edu/wp-content/uploads/2025/06/penn_demo-300x226.jpg 300w, https://db.cis.upenn.edu/wp-content/uploads/2025/06/penn_demo-768x578.jpg 768w, https://db.cis.upenn.edu/wp-content/uploads/2025/06/penn_demo-1536x1157.jpg 1536w, https://db.cis.upenn.edu/wp-content/uploads/2025/06/penn_demo.jpg 1632w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure><div class="wp-block-media-text__content">
<p class="wp-block-paragraph"><a href="https://api.zotero.org/users/3604318/publications/items/T6TZBJTL/file/view"><strong>ScaleLLM: A technique for scalable LLM-augmented data systems</strong>. </a></p>



<p class="wp-block-paragraph">Paul Loh (University of Pennsylvania); Ashwin Alaparthi (University of Pennsylvania); Ryan Marcus (University of Pennsylvania);</p>
</div></div>



<p class="wp-block-paragraph"> </p>



<ul class="wp-block-list">
<li><strong>PY-SHARQ: A Holistic Python Library for Explaining Association Rules on Relational Data</strong><br>Hadar Ben-Efraim (Bar-Ilan University), Susan Davidson (University of Pennsylvania), Amit Somech (Bar-Ilan University)</li>
</ul>



<p class="wp-block-paragraph">We hope to see you in Berlin!</p>



<p class="wp-block-paragraph"></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Penn DB Group @ VLDB 2024</title>
		<link>https://db.cis.upenn.edu/2024/07/01/penn-db-group-vldb-2024/</link>
		
		<dc:creator><![CDATA[Zack Ives]]></dc:creator>
		<pubDate>Mon, 01 Jul 2024 16:02:23 +0000</pubDate>
				<category><![CDATA[papers]]></category>
		<guid isPermaLink="false">https://db.cis.upenn.edu/?p=439</guid>

					<description><![CDATA[This August, VLDB 2024 will be in Guangzhou, China! Penn will again be well-represented, with 3 papers in the Research Track: In Towards Full Stack Adaptivity in Permissioned Blockchains, PhD student Chenyuan Wu, postdoc alumnus Mohammad Javad Amiri (Stony Brook University), undergrad student Haoyun Qin (Class of 2025), PhD student<a class="moretag" href="https://db.cis.upenn.edu/2024/07/01/penn-db-group-vldb-2024/"> Read more</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">This August, <a href="https://vldb.org/2024/">VLDB 2024</a> will be in Guangzhou, China!</p>



<p class="wp-block-paragraph">Penn will again be well-represented, with 3 papers in the Research Track:</p>


<div class="wp-block-image">
<figure class="aligncenter size-full"><img loading="lazy" decoding="async" width="658" height="330" src="https://db.cis.upenn.edu/wp-content/uploads/2024/07/image-1.png" alt="" class="wp-image-442" srcset="https://db.cis.upenn.edu/wp-content/uploads/2024/07/image-1.png 658w, https://db.cis.upenn.edu/wp-content/uploads/2024/07/image-1-300x150.png 300w" sizes="auto, (max-width: 658px) 100vw, 658px" /></figure>
</div>


<p class="wp-block-paragraph">In <strong><a href="https://dl.acm.org/doi/pdf/10.14778/3641204.3641216">Towards Full Stack Adaptivity in Permissioned Blockchains</a></strong>, PhD student <a href="https://chenyuanwu.com/">Chenyuan Wu</a>, postdoc alumnus <a href="https://www3.cs.stonybrook.edu/~amiri/">Mohammad Javad Amiri</a> (Stony Brook University), undergrad student <a href="https://haoyunqin.com/">Haoyun Qin</a> (Class of 2025), PhD student <a href="https://www.linkedin.com/in/bmehta5">Bhavana Mehta</a>, and Profs. <a href="https://rmarcus.info/">Ryan Marcus</a> and <a href="https://boonloo.cis.upenn.edu/">Boon Thau Loo</a> study the problem of supporting a (virtual) distributed database with untrusted components &#8212; using a learning-based approach. This paper articulates a vision for a learning-based untrustworthy distributed database. We focus on permissioned blockchain systems as an emerging instance of untrustworthy distributed databases and argue that as novel smart contracts, modern hardware, and new cloud platforms arise, future-proof permissioned blockchain systems need to be designed with full-stack adaptivity in mind. At the application level, a future-proof system must adaptively learn the best-performing transaction processing paradigm and quickly adapt to new hardware and unanticipated workload changes on the fly. Likewise, the Byzantine consensus layer must dynamically adjust itself to the workloads, faulty conditions, and network configuration while maintaining compatibility with the transaction processing paradigm. At the infrastructure level, cloud providers must enable cross-layer adaptation, which identifies performance bottlenecks and possible attacks, and determines at runtime the degree of resource disaggregation that best meets application requirements. Within this vision of the future, the paper outlines several research challenges together with some preliminary approaches.</p>


<div class="wp-block-image">
<figure class="aligncenter size-full is-resized"><img loading="lazy" decoding="async" width="513" height="195" src="https://db.cis.upenn.edu/wp-content/uploads/2024/07/image.png" alt="" class="wp-image-441" style="width:500px;height:auto" srcset="https://db.cis.upenn.edu/wp-content/uploads/2024/07/image.png 513w, https://db.cis.upenn.edu/wp-content/uploads/2024/07/image-300x114.png 300w" sizes="auto, (max-width: 513px) 100vw, 513px" /></figure>
</div>


<p class="wp-block-paragraph">In <strong><a href="https://www.vldb.org/pvldb/vol17/p250-naik.pdf">Relational Query Synthesis ⋈︁ Decision Tree Learning</a></strong>, PhD student <a href="https://www.seas.upenn.edu/~asnaik/">Aaditya Naik</a>, PhD alumnus <a href="https://aalok-thakkar.github.io/">Aalok Thakkar</a> (Ashoka University), PhD student <a href="https://www.seas.upenn.edu/~steinad/">Adam Stein</a>, and Profs. <a href="https://www.cis.upenn.edu/~alur">Rajeev Alur</a> and <a href="https://www.cis.upenn.edu/~mhnaik">Mayur Naik</a> address the problem of supporting <em>synthesis</em> of SQL queries and consider its interaction with machine learning. They study the problem of synthesizing select-project-join (SPJ) queries from input-output examples. Search-based synthesis techniques are suited to synthesizing projections and joins by navigating the network of relational tables but require additional supervision for synthesizing comparison predicates. On the other hand, decision tree learning techniques are suited to synthesizing comparison predicates when the input database can be summarized as a single labelled relational table. In this paper, they adapt and interleave methods from the domains of relational query synthesis and decision tree learning, and present an end-to-end framework for synthesizing relational queries with categorical and numerical comparison predicates. Their technique guarantees the completeness of the synthesis procedure and strongly encourages minimality of the synthesized program. They present <em>Libra</em>, an implementation of this technique and evaluate it on a benchmark suite of 1,475 instances of queries over 159 databases with multiple tables. Libra solves 1,361 of these instances in an average of 59 seconds per instance. It outperforms state-of-the-art program synthesis tools <em>Scythe</em> and <em>PatSQL</em> in terms of both the running time and the quality of the synthesized programs.</p>


<div class="wp-block-image">
<figure class="aligncenter size-large"><img loading="lazy" decoding="async" width="1024" height="542" src="https://db.cis.upenn.edu/wp-content/uploads/2024/07/image-2-1024x542.png" alt="" class="wp-image-443" srcset="https://db.cis.upenn.edu/wp-content/uploads/2024/07/image-2-1024x542.png 1024w, https://db.cis.upenn.edu/wp-content/uploads/2024/07/image-2-300x159.png 300w, https://db.cis.upenn.edu/wp-content/uploads/2024/07/image-2-768x407.png 768w, https://db.cis.upenn.edu/wp-content/uploads/2024/07/image-2.png 1158w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /></figure>
</div>


<p class="wp-block-paragraph">In <strong>Searching Data Lakes for Nested and Joined Data</strong>, PhD alumnus <a href="https://yizhang.io/">Yi Zhang</a> (AWS) and undergrad alumnus <a href="https://peterbaile.github.io/">Peter Chen</a> (PhD student at MIT) and Prof. <a href="https://www.cis.upenn.edu/~zives">Zachary Ives</a> consider how to perform <em>search</em> for hierarchical (JSON, Pandas dataframe) or joined data, within a data lake of data that has been indexed in first-normal-form. Exploratory data science is driving new data management platforms that assist data scientists with common tasks, such as the integration and wrangling steps required to assemble training datasets. Such tools take the data scientists’ work-in-progress data as a search object (table or JSON), and find relevant supplementary data from an organizational data lake, which can be unioned or joined with the current data – adding instances or features. Existing data lake search tools seek to find single, relational tables at a time — to match or join with a search table. Yet many data science applications revolve around finding matches to hierarchical data, which can only be matched by creating views simultaneously joining and transforming several tables in the data lake. In this paper, they extend the <a href="https://db.cis.upenn.edu/juneau-promoting-reuse-and-retargeting-in-data-science/" data-type="page" data-id="69">Juneau data lake search system</a> to search for this broader class of matches at scale. Their contribution is a general framework for efficiently merging ranked results, leveraging novel techniques for indexing and sketching, and incorporating existing single-table search techniques and ranking functions. They experimentally validate the benefits of their methods and their broad applicability using real data from data science computational notebooks. Their results indicate that, with respect to different ranking functions, their approach can return the optimal set of views up to 4.81x faster and 43% more related compared to heuristics baselines, and increase the data domain coverage by up to 28%. As a case study to show the usability of their augmentation to data science downstream tasks, their methods can reduce the regression error by up to 6.63%, and improve the classification accuracy by up to 19.5% for ML models.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Penn DB Group @ SIGMOD &#8217;24</title>
		<link>https://db.cis.upenn.edu/2024/06/14/penn-db-group-sigmod-24/</link>
		
		<dc:creator><![CDATA[Ryan Marcus]]></dc:creator>
		<pubDate>Sat, 15 Jun 2024 01:34:02 +0000</pubDate>
				<category><![CDATA[awards]]></category>
		<category><![CDATA[events]]></category>
		<category><![CDATA[papers]]></category>
		<guid isPermaLink="false">https://db.cis.upenn.edu/?p=416</guid>

					<description><![CDATA[The Penn DB Group presented a number of papers at SIGMOD 2024, hosted in Santiago, Chile! Penn presented five papers (four in SIGMOD and one in aiDM). Ph.D. student Soonbo Han (advisor: Zachary Ives) presented his work titled &#8220;Implementation Strategies for Views over Property Graphs,&#8221; which won the best paper<a class="moretag" href="https://db.cis.upenn.edu/2024/06/14/penn-db-group-sigmod-24/"> Read more</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The Penn DB Group presented a number of papers at SIGMOD 2024, hosted in Santiago, Chile! Penn presented five papers (four in SIGMOD and one in aiDM).</p>



<figure class="wp-block-gallery has-nested-images columns-2 is-cropped wp-block-gallery-1 is-layout-flex wp-block-gallery-is-layout-flex">
<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" data-id="425" src="https://db.cis.upenn.edu/wp-content/uploads/2024/06/shan_sigmod-1-1024x576.jpg" alt="" class="wp-image-425" srcset="https://db.cis.upenn.edu/wp-content/uploads/2024/06/shan_sigmod-1-1024x576.jpg 1024w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/shan_sigmod-1-300x169.jpg 300w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/shan_sigmod-1-768x432.jpg 768w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/shan_sigmod-1-1536x864.jpg 1536w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/shan_sigmod-1.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Ph.D. student Soonbo Han presenting &#8220;Implementation Strategies for Views over Property Graphs,&#8221; the best paper award winner</figcaption></figure>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" data-id="426" src="https://db.cis.upenn.edu/wp-content/uploads/2024/06/zyi_sigmod-1-1024x576.jpg" alt="" class="wp-image-426" srcset="https://db.cis.upenn.edu/wp-content/uploads/2024/06/zyi_sigmod-1-1024x576.jpg 1024w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/zyi_sigmod-1-300x169.jpg 300w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/zyi_sigmod-1-768x432.jpg 768w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/zyi_sigmod-1-1536x864.jpg 1536w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/zyi_sigmod-1.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Ph.D. student Zixuan Yi presenting &#8220;Low Rank Approximation for Learned Query Optimization&#8221; at the aiDM workshop</figcaption></figure>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" data-id="427" src="https://db.cis.upenn.edu/wp-content/uploads/2024/06/jliang_sigmod-1-1024x576.jpg" alt="" class="wp-image-427" srcset="https://db.cis.upenn.edu/wp-content/uploads/2024/06/jliang_sigmod-1-1024x576.jpg 1024w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/jliang_sigmod-1-300x169.jpg 300w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/jliang_sigmod-1-768x432.jpg 768w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/jliang_sigmod-1-1536x864.jpg 1536w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/jliang_sigmod-1.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Ph.D. student Jiaming Liang presenting &#8220;RITA: Group Attention is All You Need for Timeseries Analytics&#8221;</figcaption></figure>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="576" data-id="428" src="https://db.cis.upenn.edu/wp-content/uploads/2024/06/ytian_sigmod-1-1024x576.jpg" alt="" class="wp-image-428" srcset="https://db.cis.upenn.edu/wp-content/uploads/2024/06/ytian_sigmod-1-1024x576.jpg 1024w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/ytian_sigmod-1-300x169.jpg 300w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/ytian_sigmod-1-768x432.jpg 768w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/ytian_sigmod-1-1536x864.jpg 1536w, https://db.cis.upenn.edu/wp-content/uploads/2024/06/ytian_sigmod-1.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Visiting Ph.D. student Yao Tian presenting &#8220;A Learned Cuckoo Filter for Approximate Membership Queries over Variable-sized Sliding Windows on Data Streams&#8221;</figcaption></figure>



<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="3015" height="1989" data-id="436" src="https://db.cis.upenn.edu/wp-content/uploads/2024/06/zives_sigmod-1.avif" alt="" class="wp-image-436"/><figcaption class="wp-element-caption">Zachary Ives presenting Peizhi Wu&#8217;s paper, &#8220;Modeling Shifting Workloads for Learned Database Components&#8221;</figcaption></figure>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="5320" height="4284" data-id="434" src="https://db.cis.upenn.edu/wp-content/uploads/2024/06/group_sigmod.avif" alt="" class="wp-image-434"/><figcaption class="wp-element-caption">Yao, Zack, and Zixuan after Zixuan&#8217;s talk</figcaption></figure>
</figure>



<p class="wp-block-paragraph">Ph.D. student <a href="https://www.cis.upenn.edu/~soonbo/">Soonbo Han</a> (advisor: Zachary Ives) presented his work titled &#8220;<a href="https://dl.acm.org/doi/abs/10.1145/3654949">Implementation Strategies for Views over Property Graphs</a>,&#8221; which won the <strong>best paper award</strong>! Soonbo&#8217;s work shows query rewriting techniques can take advantage of semantic views over graph data, including how to index and maintain such views dynamically.</p>



<p class="wp-block-paragraph">Ph.D. student Jiaming Liang (advisor: Zachary Ives) presented his work titled &#8220;<a href="https://dl.acm.org/doi/10.1145/3639317">RITA: Group Attention is All You Need for Timeseries Analytics</a>.&#8221; Jiaming&#8217;s work shows how careful grouping and caching of semantically-similar inputs can accelerate neural attention mechanisms, allowing attention networks to scale up to previously-impossible tasks.</p>



<p class="wp-block-paragraph">Due to visa issues, Zack presented a paper from Ph.D. student <a href="https://www.cis.upenn.edu/~pagewu/">Peizhi Wu</a>, titled &#8220;<a href="https://dl.acm.org/doi/abs/10.1145/3639293">Modeling Shifting Workloads for Learned Database Systems</a>.&#8221; Peizhi&#8217;s work shows how to keep learned database components up to date with data drift using a carefully-tuned replay buffer.</p>



<p class="wp-block-paragraph">Visiting Ph.D. student <a href="https://www.ustyaotian.com/">Yao Tian</a> (advisor:  Xiaofang Zhou, Penn supervisors: Zachary Ives and Ryan Marcus) presented her work titled &#8220;<a href="https://dl.acm.org/doi/10.1145/3626758">A Learned Cuckoo Filter for Approximate Membership Queries over Variable-sized Sliding Windows on Data Streams</a>.&#8221; Yao&#8217;s work combines traditional Cuckoo filters with deep learning models to enable approximate membership queries over dynamically-sized windows, achieving significantly higher accuracy than previous results.</p>



<p class="wp-block-paragraph">At the aiDM workshop, Ph.D. student <a href="https://zixy17.github.io/">Zixuan Yi</a> (advisor: Ryan Marcus and Zachary Ives) presented her work titled &#8220;<a href="http://rm.cab/limeqo">Low Rank Approximation for Learned Query Optimization</a>.&#8221; Zixuan&#8217;s work shows how linear methods for approximating low rank matrices can be used to learn to steer an entire query workload at once.</p>



<p class="wp-block-paragraph">Later this summer, Penn will present several papers at VLDB in Guangzhou, China!</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Penn DB Group @ VLDB &#8217;23</title>
		<link>https://db.cis.upenn.edu/2023/09/15/penn-db-group-vldb-23/</link>
		
		<dc:creator><![CDATA[Ryan Marcus]]></dc:creator>
		<pubDate>Fri, 15 Sep 2023 19:00:00 +0000</pubDate>
				<category><![CDATA[papers]]></category>
		<guid isPermaLink="false">https://db.cis.upenn.edu/?p=335</guid>

					<description><![CDATA[The Penn Database Group recently returned from VLDB in lovely Vancouver! Penn students, faculty, and collaborators presented three research papers and two workshop papers at the conference. Ph.D. student Chenyuan Wu (supervisor: Boon Thau Loo) presented work on FlexChain and AdaChain. FlexChain (PDF) is the first work to present a<a class="moretag" href="https://db.cis.upenn.edu/2023/09/15/penn-db-group-vldb-23/"> Read more</a>]]></description>
										<content:encoded><![CDATA[
<p class="wp-block-paragraph">The Penn Database Group recently returned from VLDB in lovely Vancouver! Penn students, faculty, and collaborators presented three research papers and two workshop papers at the conference.</p>



<p class="wp-block-paragraph">Ph.D. student <a href="http://chenyuanwu.com/">Chenyuan Wu</a> (supervisor: <a href="https://boonloo.cis.upenn.edu/">Boon Thau Loo</a>) presented work on FlexChain and AdaChain. <strong>FlexChain</strong> (<a href="https://dl.acm.org/doi/pdf/10.14778/3561261.3561264">PDF</a>) is the first work to present a blockchain that disaggregates compute, storage, RAM, and cold storage to better serve diverse workloads. <strong>AdaChain</strong> (<a href="https://api.zotero.org/users/3604318/publications/items/6V8J2EL3/file/view">PDF</a>) shows how reinforcement learning can be used to adaptively switch a blockchain system between different architectures, keeping up with workload or hardware changes.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="768" src="https://db.cis.upenn.edu/wp-content/uploads/2023/09/wchen_posters-1024x768.jpg" alt="" class="wp-image-336" srcset="https://db.cis.upenn.edu/wp-content/uploads/2023/09/wchen_posters-1024x768.jpg 1024w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/wchen_posters-300x225.jpg 300w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/wchen_posters-768x576.jpg 768w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/wchen_posters-1536x1152.jpg 1536w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/wchen_posters.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Chenyuan Wu with posters for FlexChain and AdaChain.</figcaption></figure>



<div class="wp-block-group is-nowrap is-layout-flex wp-container-core-group-is-layout-7387b849 wp-block-group-is-layout-flex">
<figure class="wp-block-image size-full"><img loading="lazy" decoding="async" width="844" height="709" src="https://db.cis.upenn.edu/wp-content/uploads/2023/09/Screenshot-from-2023-09-11-12-12-20.png" alt="System architecture of AutoSteer (copied from PDF)" class="wp-image-337" srcset="https://db.cis.upenn.edu/wp-content/uploads/2023/09/Screenshot-from-2023-09-11-12-12-20.png 844w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/Screenshot-from-2023-09-11-12-12-20-300x252.png 300w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/Screenshot-from-2023-09-11-12-12-20-768x645.png 768w" sizes="auto, (max-width: 844px) 100vw, 844px" /></figure>



<p class="wp-block-paragraph"><strong>AutoSteer</strong> (<a href="https://api.zotero.org/users/3604318/publications/items/K2NZ32C4/file/view">PDF</a>), an extensible system for steering the query optimizer of any SQL database, was presented by collaborator and TUM Ph.D. student <a href="https://db.in.tum.de/~anneser/index.shtml?lang=en">Christoph Anneser</a>. AutoSteer is <a href="https://github.com/IntelLabs/Auto-Steer">open source</a>, and has been tested on Meta&#8217;s Presto cluster. The AutoSteer work was a collaboration between UPenn, Intel Labs, MIT, Meta, and TU Munich. </p>



<p class="wp-block-paragraph"></p>



<p class="wp-block-paragraph"></p>
</div>



<p class="wp-block-paragraph">Ph.D. student Bhavana Mehta (supervisor: <a href="https://boonloo.cis.upenn.edu/">Boon Thau Loo</a>) presented <strong>RLShard</strong> (<a href="https://api.zotero.org/users/3604318/publications/items/AAG2JG9F/file/view">PDF</a>), a vision for a Byzantine fault tolerant, sharded, transactional database at the aiDB workshop. RLShard contains designs for both a fully decentralized, trustless system, as well as a semi-centralized version taking advantage of an administrative domain. Bhavana will continue developing both approaches as part of her dissertation.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="336" src="https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_220612202.MP3-1024x336.jpg" alt="" class="wp-image-338" srcset="https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_220612202.MP3-1024x336.jpg 1024w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_220612202.MP3-300x98.jpg 300w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_220612202.MP3-768x252.jpg 768w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_220612202.MP3-1536x503.jpg 1536w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_220612202.MP3-2048x671.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Bhavana presenting RLShard</figcaption></figure>



<p class="wp-block-paragraph">Also at aiDB, assistant professor <a href="https://ryanmarc.us">Ryan Marcus</a> presented a vision for <strong>Learned Query Superoptimization</strong> (<a href="https://api.zotero.org/users/3604318/publications/items/6NQN3S6F/file/view">PDF</a>), a set of new techniques to improve common repetitive queries in analytic database systems. Ryan plans to implement this vision in the coming years.</p>



<figure class="wp-block-image size-large"><img loading="lazy" decoding="async" width="1024" height="285" src="https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_2141396301-1024x285.jpg" alt="" class="wp-image-339" srcset="https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_2141396301-1024x285.jpg 1024w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_2141396301-300x83.jpg 300w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_2141396301-768x214.jpg 768w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_2141396301-1536x427.jpg 1536w, https://db.cis.upenn.edu/wp-content/uploads/2023/09/PXL_20230901_2141396301-2048x570.jpg 2048w" sizes="auto, (max-width: 1024px) 100vw, 1024px" /><figcaption class="wp-element-caption">Ryan presenting learned query superoptimization</figcaption></figure>



<p class="wp-block-paragraph">The Penn DB Group will see you at the next one!</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
