Andy Pavlo 加入 ClickHouse 并成立 ClickHouse Labs
Andy Pavlo joins ClickHouse to establish ClickHouse Labs

原始链接: https://clickhouse.com/blog/andy-pavlo-joins-clickhouse

前卡内基梅隆大学教授兼数据库专家将加入 ClickHouse,创立并领导一个新的研究机构“ClickHouse Labs”。基于他对 ClickHouse 高性能 C++ 架构的长期推崇,他旨在打造一个行业领先的研究中心,以弥合学术理论与实际应用之间的鸿沟。 与独立的研究小组不同,ClickHouse Labs 将直接与工程团队整合,以改进现有技术并加速实验性优化的部署。该实验室将同时关注 ClickHouse 和 PostgreSQL,致力于解决性能、可靠性以及人工智能技术集成方面的挑战。具体而言,团队将探索数据库如何更好地支持自主智能体,以及这些智能体如何反过来实现数据库开发的自动化与改进。 通过将科学严谨性与工程实践相结合,ClickHouse Labs 期望能够重现 IBM 和微软研究院等传奇行业研究机构的影响力。这一举措展现了该机构在推动数据库系统基础科学发展的同时,确保 ClickHouse 始终处于行业前沿的决心。

数据库专家兼教育家 Andy Pavlo 已加入 ClickHouse,并将负责领导新成立的研究机构“ClickHouse Labs”,致力于推动数据库技术的发展。该计划旨在弥合学术研究与工程实践之间的鸿沟,其模式与 IBM Research 和 Microsoft Research 等将行业与研究相结合的模式相呼应。 这一公告在 Hacker News 上引发了热烈讨论。许多用户对这一举措表示赞赏,称赞 Pavlo 在卡内基梅隆大学(CMU)极具影响力的数据库系列课程,以及他长期以来对 SQL 的倡导。支持者认为,该实验室将成为一个至关重要的桥梁,使 ClickHouse 能够保持其在以性能为核心的分析型数据库领域“处于最前沿”的地位。 批评者和质疑者则对商业产品设立专门研究机构的必要性提出了疑问,一些人还就“硬科技”的定义展开了辩论,并讨论大学环境是否仍是进行高影响力数据库创新的最佳场所。尽管各方反应不一,但普遍共识是:Pavlo 的加入是 ClickHouse 生态系统的一个重要里程碑,标志着公司将在突破数据库性能极限方面投入巨资。
相关文章

原文

I am excited to announce that I am joining ClickHouse to establish and lead a new research team called ClickHouse Labs. I want to share how it came about and what we plan to do.

I started as a professor in the Computer Science Department at Carnegie Mellon University in 2013. I have spent my career seeking to understand the science of modern database management system (DBMS) internals. I make it a priority to track every new system that comes along, both in industry and academia, to understand their implementations.

I have known about the ClickHouse DBMS since it was first announced as open-source software in June 2016. My initial reaction to this news was that it had to be vaporware because it seemed too good to be true. ClickHouse had features that at the time were only found in a handful of closed-source, commercial analytical DBMSs. For example, ClickHouse was written in C++ and supported vectorized query execution using SIMD in 2016. Most prominent open-source analytical DBMSs in 2016 were JVM-based and did not support SIMD optimizations until years later.

Since then, I have followed ClickHouse's development closely. It has always been a leading system that was highly relevant to our academic research projects. You can even see me wearing my original ClickHouse shirt in my first remote lectures in 2020, when the pandemic forced us to move our database courses online.

Given this history, I was honored when the ClickHouse co-founders invited me to establish this new research group at ClickHouse. The chance to work with one of the strongest engineering teams on the next generation of database technology was an opportunity that I could not pass up. This will be a next-level collaboration like when Killer Mike hooked up with El-P to create a hip-hop supergroup.

The goal of ClickHouse Labs is to establish a best-in-class industry research organization focused on databases. It will not operate as an isolated research organization that throws ideas over the wall to engineering. Instead, we will work closely with ClickHouse engineers, customers, collaborators, and industry partners to develop and disseminate new ideas that keep ClickHouse at the bleeding edge.

We will also work with ClickHouse's PostgreSQL team to help establish its burgeoning managed service as a market leader in performance and reliability. PostgreSQL and ClickHouse serve different workload requirements, but the combination gives us a broad foundation for investigating both transactional and analytical database problems.

Our objective is straightforward but ambitious: conduct research with scientific value and then help transform the best ideas into technology that matters to users. I want to achieve the same level of impact associated with pioneering industry research organizations, such as IBM Research and Microsoft Research. Those groups demonstrated that industry laboratories can simultaneously advance fundamental computer science, influence commercial products, and train generations of database researchers. That is the tradition we want to continue.

The ClickHouse team already has an exceptional record of publishing deep technical material about its work. Since the establishment of the company in 2021, its engineers have produced detailed articles that explain the DBMS's implementation. There is also the 2024 VLDB paper that describes ClickHouse's core architecture. These works are so thorough that I assign them as readings to my students at Carnegie Mellon. At the same time, there is a backlog of interesting ideas and optimizations that the ClickHouse engineering team has explored but has not yet had the time to validate fully and push into production. One of my immediate priorities is to help accelerate this process. We will then use that as a springboard to explore new ideas that push ClickHouse even further.

One larger question we will investigate is how DBMSs like ClickHouse and PostgreSQL fit into emerging AI and agentic technologies. There are two sides to this problem. The first is determining what a DBMS should look like to better support agents. The second side is determining how agents can improve and automate the development of DBMSs themselves. Everything is on the table: new hardware, new algorithms, new data structures, new execution strategies, and new ways of building and operating DBMS software. Although I do not have answers to these problems yet (this is why it is research), the one thing I am certain about is that ClickHouse's solid relational model foundation positions it well to evolve alongside these data-intensive workloads.

I have spent my career studying how database systems are built and helping train the people who build them. With ClickHouse Labs, we now have the opportunity to create an organization devoted to advancing both.

联系我们 contact @ memedata.com