【AI前沿】CaaS崛起:Agentic AI的上下文即服务
字幕摘录
| 时间 | 英文 | 中文 |
|---|---|---|
| 0:13 | So, hi, everyone. | 大家好,你们好。 |
| 0:14 | Thank you so much for taking the time to join this session. | 非常感谢你们抽出时间参加本届会议。 |
| 0:16 | I hope I'll, or at least I can guarantee you, I'll do whatever it takes to make it worth | 我希望我会,或至少我可以保证, 我会不惜一切代价使它变得值得 |
| 0:20 | your time. | 你的时间。 |
| 0:22 | My name is Omel. | 我叫奥梅尔 |
| 0:23 | I lead the product marketing team over at Bright Data. | 我在Bright Data公司领导产品营销小组 |
| 0:26 | Just by maybe a quick show of hands. | 也许只是举手 |
| 0:27 | Who here is familiar with Bright Data? | 谁在这里熟悉亮数据? |
| 0:31 | Okay. | 摆 |
| 0:31 | We can do better. | 我们可以做得更好。 |
展开字幕全文(386 条)
| 序号 | 英文 | 中文 |
|---|---|---|
| 1 | So, hi, everyone. | 大家好,你们好。 |
| 2 | Thank you so much for taking the time to join this session. | 非常感谢你们抽出时间参加本届会议。 |
| 3 | I hope I'll, or at least I can guarantee you, I'll do whatever it takes to make it worth | 我希望我会,或至少我可以保证, 我会不惜一切代价使它变得值得 |
| 4 | your time. | 你的时间。 |
| 5 | My name is Omel. | 我叫奥梅尔 |
| 6 | I lead the product marketing team over at Bright Data. | 我在Bright Data公司领导产品营销小组 |
| 7 | Just by maybe a quick show of hands. | 也许只是举手 |
| 8 | Who here is familiar with Bright Data? | 谁在这里熟悉亮数据? |
| 9 | Okay. | 摆 |
| 10 | We can do better. | 我们可以做得更好。 |
| 11 | I'll pass it on to our brand team. | 我会把它传给我们的品牌团队。 |
| 12 | Bright Data是一家网络数据公司. | |
| 13 | Basically we help more than 20,000 teams around the world, including more than 70% of the | 基本上我们帮助了全球两万多支球队,包括超过70%的 |
| 14 | world's biggest AI labs, to extract data from the web. | 世界上最大的人工智能实验室 从网上提取数据 |
| 15 | Just to put this in perspective of what scale we're talking about, we're talking well over | 仅仅从我们所说的规模的角度来说,我们谈得很好 |
| 16 | 50 billion pages, HTMLs every day, more than 20 petabytes of video, audio, and other media | 500亿页, 每天 HTML, 超过 20 petabytes 的视频,音频,和其他媒体 |
| 17 | data. | 数据。 |
| 18 | So that's just the perspective of the type of work that we do at Bright Data. | 这就是我们在Bright Data公司的工作类型。 |
| 19 | But enough about us. | 但是关于我们已经够了。 |
| 20 | Personally, I joined Bright Data about three years ago, which essentially gave me front | 就我个人而言,我三年前加入了 Bright Data, 这基本上让我站在前面 |
| 21 | row seats at everything around AI and the web and how they started to actually connect. | 在AI和网络周围的所有位置排成一排, 以及它们是如何开始连接的。 |
| 22 | It sounds very old, but if you think about it, only maybe less than two years ago, right, | 听起来很老了,但是如果你考虑一下, 也许不到两年前,对吗, |
| 23 | we were able to start using, access the web, and search the web through cloud or through | 我们得以开始使用、访问网络并通过云或通过网络搜索网络 |
| 24 | 闲聊GPT. | |
| 25 | That option didn't even exist in the earlier versions, right? | 这个选项在早期版本中甚至不存在,对吧? |
| 26 | So that connection, that way in which both AI and the web are starting to converge is | 因此,这种连接, AI和网络开始融合的方式是 |
| 27 | something that is still evolving and evolving rapidly. | 一些仍在迅速演变中的东西。 |
| 28 | That's part of what I want to try and shed some light on today and talk about a new emerging | 这也是我想尝试的一部分 并给今天一些启示 并谈论一个新的新兴 |
| 29 | breed of companies that's coming out of this connection. | 由这种联系产生的公司 |
| 30 | I think we can basically agree that the web is by far the world's greatest source of data. | 我认为我们基本上可以同意, 网络是世界上最大的数据来源。 |
| 31 | At least historically, when it comes to Bright Data, that's all we cared about, right, helping | 至少在历史上,当它涉及到亮数据, 这就是我们关心的,正确的,帮助 |
| 32 | our customers extract data from the web. | 我们的客户从网上提取数据 |
| 33 | But with the emergence of AI and the emergence more recently of AI agents that need to do | 但随着AI的出现 和最近出现的AI代理 需要做 |
| 34 | knowledge work, right, the web is no longer just a source of data. | 知识工作,对了,网络不再仅仅是数据的来源. |
| 35 | We can actually start looking at it as a source of context. | 我们实际上可以把它看成是背景的来源。 |
| 36 | Context in the sense that if I do knowledge work and I have knowledge agents that support | 也就是说,如果我做知识工作,我拥有支持的知识代理人 |
| 37 | my work, I want to go out to the web, find the information that I need, use it as context, | 我的工作,我想去网络, 找到我需要的信息,用它作为背景, |
| 38 | but keep on working. | 但继续工作 |
| 39 | So the data itself is only a step in the process for something bigger, for the actions I need | 因此,数据本身只是进程的一个步骤 对于更大的东西, 对于我需要的行动 |
| 40 | to take, for the conclusions I need to draw, and for every downstream application that | 对于我需要得出的结论,对于每一个下游应用 |
| 41 | follows. | 接下来。 |
| 42 | The first one is to figure this one out, oh, sorry, even before that. | 第一个是想出这个, 哦,对不起,甚至在那之前。 |
| 43 | But one thing that I want all of us to bear in mind because this is going to follow us | 但有件事我希望大家记住 因为这会跟着我们 |
| 44 | through the rest of this conversation. | 通过其余的谈话。 |
| 45 | The web is messy, it's unstructured, and most importantly, it changes all the time. | 网络杂乱无章,结构不整齐,最重要的是,经常变化. |
| 46 | This is a chart that shows data decay, right, that it's analysis done by our team that shows | 这是一张显示数据衰减的图表,没错, 它是我们团队所做的分析显示 |
| 47 | data decay basically how long after a new page, a new piece of content goes live, it | 数据衰减 基本上在一个新页面之后多久 一个新的内容开始运行,它 |
| 48 | is no longer relevant, right? | 已经无关紧要了吧? |
| 49 | So social media, it's easy for us to understand, it's far less than a day. | 所以社交媒体,我们很容易理解,还不到一天. |
| 50 | But also news, finance, retail, 30 days later, data that was collected, it's mostly no longer | 但新闻,金融,零售,30天后, 收集到的数据, 大部分已经不再 |
| 51 | relevant. | 相关。 |
| 52 | And when we acknowledge that, this simple notion, we understand that extracting context | 而当我们承认,这个简单的概念, 我们理解提取的背景 |
| 53 | from the web or relying on the web is not a snapshot, it's not a one-time effort, it's | 从网络或依赖网络不是快照,不是一次性努力,而是 |
| 54 | not even a monthly effort. | 连月费都没有 |
| 55 | It's something that we need to keep on doing. | 这是我们需要继续做的事情。 |
| 56 | It's something we need to look at as an ongoing process and something that we need to be mindful | 我们需要将它视为一个持续的过程 和我们需要注意的东西 |
| 57 | The first ones to figure it out were, of course, search companies, right? | 最早发现的当然是搜索公司,对吧? |
| 58 | Only three years ago, we were on the far left, right? | 仅仅三年前,我们在最左边,对不对? |
| 59 | This is in our lifetime, right? | 这是在我们一生中,对不对? |
| 60 | Three years ago, we were on the far left, everything was Google. | 三年前,我们在远方, 一切都是谷歌。 |
| 61 | There was complete and total dominance up until three years ago for the past 20 or so | 过去20年左右的三年前 完全和完全的主导地位 |
| 62 | years. | 岁月 |
| 63 | That's what we're talking about. | 师便打曰. |
| 64 | Purely for humans, search something, go, collect the information you need, and carry on. | 纯粹为人类,搜索一些东西,去,收集你需要的信息,然后继续. |
| 65 | Then, fast forward, maybe one, half years ago, two years ago, search began to appear | 然后,快速前进,也许一年半前,两年前, 搜索开始出现 |
| 66 | within the LLMs, within the chatbots, right? | 在LLMs内部, 在聊天室内部,对不对? |
| 67 | Which already started blurring the line between humans and agents because now the same bots | 已经模糊了人类和特工之间的界限 因为现在都一样 |
| 68 | also have the same web search available, so the same LLMs have the same access through | 也可以使用同样的网页搜索,所以同样的LLMs通过 |
| 69 | API for the bots, so we started seeing that convergence happening, right? | 蛋蛋的API, 所以我们开始看到 交汇的发生,对不对? |
| 70 | And so, for the first time, we're seeing more and more traffic flowing down, search traffic, | 所以,我们第一次看到 越来越多的流量流下,搜索流量, |
| 71 | search intent flowing down these channels, not only to Google. | 搜索意图流出这些频道, 不仅仅是谷歌。 |
| 72 | And last but not least, right, we now see a whole breed of companies, the AI search | 最后但并非最不重要的是,没错,我们现在看到的是 整个品种的公司,AI搜索 |
| 73 | companies, which you're familiar with them. | 公司,你熟悉它们。 |
| 74 | I caught a talk yesterday by Will, the CEO of Exxon, and Parallel, and U.com, and Tavili, | 我昨天接到了埃克森公司总裁威尔 和平行公司和U.com公司 以及塔维利的谈话 |
| 75 | and a bunch of others. | 还有一群人 |
| 76 | They are purely built in indexing the web, especially for agents, not even looking at | 它们纯粹是用来编制网络索引的 特别是针对代理商 连看都不看 |
| 77 | the humans involved anymore. | 与人类有关 |
| 78 | So Google's dominance when it comes to search, if Google was synonymous of web search, that | 所以谷歌在搜索时的主导地位 如果谷歌是网络搜索的同义词的话 |
| 79 | is very much shaking. | 非常颤抖。 |
| 80 | When there's blood in the water, the sharks come. | 当水中有血,鲨鱼就来了. |
| 81 | Just last week, Amazon, Amazon announced, I don't know how many of you saw it, that | 就在上周,亚马逊,亚马逊宣布, 我不知道有多少人看到它, |
| 82 | they developed their own index and started allowing the ability to retrieve data from | 他们开发了自己的索引,并开始允许从中检索数据的能力。 |
| 83 | the web, to retrieve context for agents on agent core. | 网络,以检索代理核心上的代理上下文。 |
| 84 | Amazon developed their own search engine. | 亚马逊开发了自己的搜索引擎. |
| 85 | Two weeks before that, it was Microsoft. | 两周前是微软公司 |
| 86 | Microsoft always had skin in the game. | 微软在游戏中总是有皮肤. |
| 87 | I don't know, one or two percent of the world's search traffic went to Microsoft. | 我不知道,全球搜索流量的一二成都投给了微软. |
| 88 | But they have now repackaged it and launched it, again, as part of Web IQ, as part of their | 但他们现在重新包装 并再次推出它 作为Web IQ的一部分 |
| 89 | suite for agentic development and orchestration. | 用于代理开发和管弦乐的套房。 |
| 90 | So we're seeing more and more this space becoming crowded. | 我们看到这个空间越来越拥挤。 |
| 91 | But when we're talking about context, and we're looking at this through the lens of | 但当我们谈论背景时, 我们从镜头看这个 |
| 92 | search, I believe it only tells us part of the story, right? | 搜索,我认为它只告诉我们 故事的一部分,对不对? |
| 93 | I can search for, I don't know, what's the cost of a certain pair of sneakers this morning, | 我可以搜索,我不知道, 什么是成本 某双运动鞋今天上午, |
| 94 | right, on a certain website. | 对,在某个网站上 |
| 95 | I cannot really search for how has that price changed over the last six months, what discount | 我实在找不到过去六个月价格的变动 |
| 96 | it had, right? | 曾经是 对吧? |
| 97 | 我能找到Bright Data的公开职位 | |
| 98 | We do. | 没错 |
| 99 | I urge you to go have a look. | 我劝你去看看 |
| 100 | But I can't see how that was a chart and how that changed over time and how the headcount | 但我看不出这是一张图表 和它如何变化 随着时间和人口统计 |
| 101 | of the company changed over time. | 公司随时间而变化。 |
| 102 | And all of that information existed on the web simply back then. | 所有的信息都存在于网络上 就在那时。 |
| 103 | So when we actually start to think about it, we understand that there's much more context | 所以,当我们真正开始思考的时候, 我们明白,还有更多的背景 |
| 104 | in the web than what web search allows us to extract. | 在网络中比什么网络搜索让我们可以提取。 |
| 105 | And this is what we started seeing in the recent years, a whole new breed of companies | 这是我们近年来开始看到的 全新的公司 |
| 106 | rising. | 上升。 |
| 107 | We like to call them internally CAS, context as a service, because that's what they do. | 我们喜欢在内部称之为CAS, 上下文作为一种服务, 因为他们这样做。 |
| 108 | They allow agents to tap into them, MCP, CLI, just pure good old API, and actually start | 他们允许特工挖掘他们,MCP,CLI, 只是纯粹好的 API,实际上开始 |
| 109 | extracting data to retrieve data so they can reason over for whatever knowledge work they | 提取数据以检索数据,以便他们能够为任何知识工作提出理由 |
| 110 | are responsible for. | 负责。 |
| 111 | We see this happening in e-commerce. | 我们在电子商务中看到这种情况。 |
| 112 | We see this happening in travel. | 我们在旅行中看到这种情况。 |
| 113 | We see this happening in finance, in market research, in HR, in real estate, in a bunch | 我们从金融、市场研究、人力资源、房地产、一堆 |
| 114 | of other domains. | 其它领域。 |
| 115 | I'll show a few examples in a second, right? | 我马上会举几个例子,对吗? |
| 116 | What all of these have in common is that they don't only just discover the web, you know, | 这些东西的共同点是 它们不只是发现网络 |
| 117 | in terms of think crawling, think searching, think all of that, accessing, extracting the | 思考爬行,思考搜索, 思考这一切, 访问,提取 |
| 118 | data and indexing it. | 数据和索引。 |
| 119 | They take it a step further. | 他们更进一步。 |
| 120 | They actually develop knowledge graphs to start structuring all of the entities and | 他们实际上开发了知识图 开始构建所有实体和 |
| 121 | to dedupe them. | 驱除他们。 |
| 122 | And they start enriching them with a lot of different sources. | 他们开始用许多不同的来源来丰富它们。 |
| 123 | So they actually start merging all of that data. | 所以他们开始合并所有的数据。 |
| 124 | If you think about it, they kind of behave like vertical search engines, right? | 想想看,他们的行为就像垂直搜索引擎,对吧? |
| 125 | They are a very, very, very good search engine for something very specific. | 它们是一个非常,非常,非常好的搜索引擎 对于非常具体的东西。 |
| 126 | And it's already in full motion, right? | 并且它已经完全运动了,对不对? |
| 127 | So as I said, we see this in finance, in market research, in retail, in e-commerce, in GTM, | 因此,正如我所说,我们看到这一点在金融,市场研究,零售,电子商务,GTM, |
| 128 | in sales intelligence. | 在销售情报。 |
| 129 | What all of these companies, by the way, have in common, they're all part of Bright Data's | 顺便说一句,所有这些公司的共同点 都是Bright Data公司的一部分 |
| 130 | startup program if you are a builder. | 如果您是构建者, 则启动程序 。 |
| 131 | And this is a hot space to go in because I think we're only tapping the surface. | 这是一个热的空间 进入,因为我认为我们只是 敲打表面。 |
| 132 | I invite you to scan this and apply up to $20,000 in credits and all sorts of co-marketing. | 我邀请你扫描一下 并申请最高2万元的信用和各种共同营销。 |
| 133 | But that's enough self-promotion. | 但是,这已经足够自我促进了. |
| 134 | So CAS as an industry is already in full bloom. | 所以CAS作为一个行业已经完全开花。 |
| 135 | And as always with these situations, right, also the traditional players aren't left too | 和往常一样 传统球员也不剩了 |
| 136 | much behind. | 相当落后 |
| 137 | There at the bottom, you see good old data as a service, you see ZoomInfo, right? | 在底部,你看到好的老数据 作为一种服务,你看到ZomoInfo,对不对? |
| 138 | By researching for this presentation today, I also saw that they launched that thing at | 通过研究今天的演讲,我还看到,他们推出的东西在 |
| 139 | the top. | 顶部。 |
| 140 | 其名称为GTM.AI. | |
| 141 | You can only imagine how much they paid for that domain. | 你只能想象他们为此付出了多少钱 |
| 142 | But they launched a secondary brand for ZoomInfo that is catering specifically for the need | 但是他们推出了ZoomInfo的二级品牌 专门满足需求 |
| 143 | of agents. | 特工人员。 |
| 144 | Look at the wording, right? | 看看措辞,对不对? |
| 145 | They talk about GTM work, right, that knowledge work, that research that you do when you need | 他们谈论GTM的工作,对,知识的工作, 研究,当你需要的时候 |
| 146 | to prospect, when you need to do headhunting, whatever it is you need to do that involves | 当你需要猎头的时候 不管你需要做什么 |
| 147 | people mostly, straight from cloud code, straight from codecs or any other agent. | 大部分人,直接从云码, 直接从解码器或任何其他剂。 |
| 148 | They understand the gap, right? | 他们理解差距,对不对? |
| 149 | So yeah, it's fine to think of CAS as an evolution of DAS, and it is in a way, but it's catering | 所以,这是很好的认为 CAS 是DAS的进化, 它在某种程度上,但它的餐饮 |
| 150 | 与Google不同。 | |
| 151 | When agents need them, it's different than people. | 当特工需要他们时,它和人不同. |
| 152 | When we let this one sink, that at the very least we have two different types of paths | 当我们让这个沉下去的时候 至少我们有两种不同的路径 |
| 153 | to complete knowledge work as an agent, we can start thinking about this in terms of | 作为代理完成知识工作,我们可以开始思考这个问题 |
| 154 | web context engineering. | 网络上下文工程. |
| 155 | We can start thinking about this in terms of how do I optimize for the specific task, | 我们可以从我如何优化具体任务的角度来思考这个问题, |
| 156 | more importantly when things come as it is, how do I optimize this for breeds of tasks? | 更重要的是,当事情到来,它是什么, 我如何优化这个品种的任务? |
| 157 | How do I do this for various parts of the organization that I'm building for? | 我如何为我建立的组织的各个部分做这个? |
| 158 | If I'm an AI engineer, I need to serve different teams, they may have different needs. | 如果我是AI工程师 我需要为不同的团队服务 他们可能有不同的需要 |
| 159 | It's very tempting to throw AI search at all of them, but maybe that's not optimal. | 将AI搜索扔到他们身上是很诱人的,但也许这不是最佳的. |
| 160 | Maybe I need a combination of both. | 也许我需要两者的组合 |
| 161 | Maybe I can start seeing all sorts of cost efficiencies emerge from that. | 也许我可以开始看到 各种成本效率的出现。 |
| 162 | So for the second half of this presentation, we actually went ahead and created a test. | 因此,对于这个介绍的后半部分,我们实际上做了一个测试。 |
| 163 | This is not a benchmark. | 这不是一个基准。 |
| 164 | You won't see any something concrete that I can say with great confidence other than | 你将看不到任何具体的东西 我可以非常自信地说 |
| 165 | the actual research that we did because we wanted to start unraveling the different considerations | 我们所做的实际研究 因为我们想要 开始解开不同的考虑 |
| 166 | and how do these two stack up against each other. | 这俩人怎么互相对抗 |
| 167 | So we designed a test. | 所以我们设计了一个测试。 |
| 168 | We went for something basic. | 我们去找一些基本的东西。 |
| 169 | We said, okay, let's take a company, an entity, and try and enrich it across 25 different | 我们说,好吧,让我们采取一个公司,一个实体, 并尝试并丰富它 超过25个不同的 |
| 170 | fields. | 字段。 |
| 171 | Some of them are very easy, you know, the company domain, the name, the headquarters, | 其中一些非常简单,你知道,公司域名,名称,总部, |
| 172 | but some are more challenging, right, things about hiring and people and something. | 但有些更具有挑战性,对, 有关雇用和人的东西。 |
| 173 | And we built a simple agent, a loop that uses Opus 4.8 as the harness, and it starts | 我们建造了一个简单的代理,一个循环 使用Opus 4.8作为绳索,它开始 |
| 174 | to go field by field, go out, search for it, or retrieve it from the CAS, do it again and | 逐个出场,出去搜索,或者从CAS中取回,再做一次 |
| 175 | again and again until it completes and brings back, sets some guardrails, you know, like | 一次又一次 直到它完成和带回来, 设置一些护栏,你知道,喜欢 |
| 176 | budget and stuff just to keep it fair. | 预算什么的 只是为了保持公平 |
| 177 | And I'm going to share with you the results. | 与你分享结果 |
| 178 | So the first thing that we would care about, right, being knowledge work would be, sorry, | 所以,第一件事,我们会关心, 对,作为知识工作是,对不起, |
| 179 | we ran it 100 times on all of the sponsors of today's event. | 今天所有赞助商上百次 |
| 180 | So the first thing that we saw in terms of coverage is that there's pretty good convergence. | 因此,我们首先看到的是, 在覆盖面方面, 存在着相当好的趋同。 |
| 181 | They all did fairly well, right? | 他们都做得很好,对不对? |
| 182 | I'll get to the two at the bottom in a second. | 我马上到底部两个 |
| 183 | So search were consistent performance, one of the major CAS providers were also very | 所以搜索是一致的, 一个主要的CAS提供商 也非常 |
| 184 | well. | 不错 |
| 185 | The third one, by the way, you can see on Locker and SERP, SERP is good old data, good | 第三个,顺便说一句,你可以看到 在洛克和SERP,SERP是好的旧数据,好的 |
| 186 | old Google. | 旧谷歌. |
| 187 | We basically did the same thing just with Google, and it performed pretty well in extracting | 我们基本上做了同样的事情 与谷歌, 它的表现相当不错 提取 |
| 188 | that information. | 那个信息 |
| 189 | Native is Claude's own search, and you see that they converge really well. | 原生地是克劳德自己的搜索, 你可以看到,它们真的汇合得很好。 |
| 190 | I was originally surprised about the two CAS solutions at the bottom. | 我最初对下面的两个CAS解决方案感到惊讶. |
| 191 | It was counterintuitive. | 这是反直觉的。 |
| 192 | I expected CAS to dominate this thing because that's, you know, you had one job, right, | 我期望CAS能主导这件事 因为你有一份工作 |
| 193 | to map out these companies. | 以规划这些公司。 |
| 194 | But after diving into it a bit more, you understand that, well, they are limited in the sense | 但潜水多一点后,你就会明白,从意义上来说,它们有限 |
| 195 | that they know what they have about an entity. | 他们知道自己对于一个实体有什么感觉。 |
| 196 | If I ask them the question that is beyond that, they will never have that data, right? | 如果我问他们一个超出这个范围的问题, 他们永远不会有这些数据,对不对? |
| 197 | Unlike a search, it can go out and continue searching and exploring it. | 与搜索不同,它可以外出继续搜索和探索. |
| 198 | If they didn't collect data about the recent job hiring, it will never be there, right? | 如果他们不收集有关最近招聘工作的数据,那就永远也不会存在,对吧? |
| 199 | So it makes sense that they are a bit behind, but I'm sure at the same time that they have | 所以说他们有点落后 但我确定他们同时 |
| 200 | a lot of other advantages that we simply didn't ask for, a lot of other fields that they didn't | 我们根本没要求的很多其他优势, 很多其他的领域,他们没有 |
| 201 | have that aren't represented. | 没有代表。 |
| 202 | So again, it creates some complexities on how do we measure coverage when it relates | 因此,它再次造成一些复杂因素,说明我们如何衡量涉及 |
| 203 | to the specific job that we need to do rather than in general. | 以完成我们所需要的具体工作,而不是一般工作。 |
| 204 | The second thing we looked at was cost, of course. | 我们看到的第二件事当然是代价。 |
| 205 | Here we started seeing it spread out a bit. | 在这里,我们开始看到它扩散了一点。 |
| 206 | You can see that massive bulk in the center. | 你可以看到中央那块大块 |
| 207 | Most of the search and the CAS and even using Google, right, converged to pretty much the | 大部分搜索和CAS,甚至使用Google,对, 汇合到几乎大部分 |
| 208 | same cost, only different, right? | 同样的成本,只是不同,对不对? |
| 209 | The CAS was just about the service itself, what you pay the vendor, right? | CAS只是关于服务本身, 你付给供应商什么,对不对? |
| 210 | All of the other search solutions, you also needed a lot of token burn to actually structure | 所有其它的搜索解决方案,你还需要很多符号燃烧来实际结构 |
| 211 | that data so you can actually act on it and use it as something retrievable, right? | 数据,这样你就可以 实际采取行动,并用它 作为可检索的东西,对不对? |
| 212 | So it's the same output. | 故同输出. |
| 213 | Native, obscenely expensive, and the CAS on the right, I'm sure you're all familiar with, | 本地人 淫秽昂贵 右边的CAS 我肯定你们都熟悉 |
| 214 | the barf are the most expensive in the industry. | 烤肉是这个行业最贵的 |
| 215 | I will not name and shame them. | 我不会给他们起名和羞辱 |
| 216 | Interesting, you see that small CAS there at the left, that CAS number two, they were | 有趣的是,你看左边那个小CAS, 第二CAS,他们是 |
| 217 | very cheap. | 非常廉价。 |
| 218 | And they're also the ones that are here at the bottom, which is funny because what I | 他们也是最底层的人 这很有趣 因为我 |
| 219 | believe is happening there is that we're seeing, even within this industry, niche players that | 我们正看到 即使在这个行业里 也有合适的角色 |
| 220 | have lower quality data but much cheaper, they're already carving that niche of the | 数据质量较低,但价格便宜得多 他们已经在刻画了 |
| 221 | long tail, right, of small shops or small usage so that they don't want to pay as much | 长尾巴,右,小商店或小用途 以免他们想支付那么多 |
| 222 | and don't need as much data. | 不需要那么多数据 |
| 223 | And we're all seeing them branch out there. | 我们都看到他们的树枝在那里。 |
| 224 | Most of you here, I presume, are engineers. | 我想你们大多数是工程师 |
| 225 | So there's a very evident question that we did not ask here, which is, what is the one | 所以有一个非常明显的问题,我们没有在这里问, 那就是,什么是 |
| 226 | thing that an engineer would care about? | 工程师会关心的? |
| 227 | Thank you. | 谢谢 |
| 228 | Let's talk about scale. | 我们来谈谈规模 |
| 229 | This is the cost, not for the whole hundred, this is the cost per one, for one record. | 这是成本,而不是整个100, 这是每个成本, 记录。 |
| 230 | What happens if we need a million? | 如果我们需要一百万怎么办? |
| 231 | Now yes, a million records will not fit in the context we're obviously, we're not talking | 现在,是的,一百万的唱片 将不符合的背景 我们显然,我们不是谈论 |
| 232 | about a single run that needs a million. | 单程需要一百万 |
| 233 | You can think about a million in terms of the frequency, right? | 你可以考虑一百万的频率,对不对? |
| 234 | If I am a market research, I do the diligence for private equity, I revisit these companies | 如果我是市场研究,我做私募股权的尽职调查,我再看看这些公司 |
| 235 | all the time. | 所有的时间。 |
| 236 | I ask more questions about them as the time goes by. | 随着时间的流逝,我问更多关于他们的问题。 |
| 237 | Was there any new news about them? | 他们有什么新消息吗? |
| 238 | Was there anything that changed? | 有什么变化吗? |
| 239 | Did somebody join? | 有人加入吗? |
| 240 | Did somebody leave? | 有人走了吗? |
| 241 | Do they have new hires? | 他们有新工作吗? |
| 242 | I keep on asking the same thing. | 我一直问同样的事情。 |
| 243 | So when I'm talking about this, the multiply by a million, it's not just about the number | 所以当我谈论这个,乘以一百万, 它不仅仅是数字 |
| 244 | of companies, it's the frequency in which I'm asking it. | 公司,这是频率 我问它。 |
| 245 | Frequency is the cost killer when we talk about these, and we need to acknowledge that, | 频率是成本杀手 当我们谈论这些, 我们需要承认, |
| 246 | right? | 对吧? |
| 247 | We're thinking about this in terms of web context engineering. | 我们在网络环境工程方面考虑这个问题。 |
| 248 | We're starting to look at it differently. | 我们开始不同看待它。 |
| 249 | Every repeated query costs the same as the first. | 每次重复查询的费用与第一次相同. |
| 250 | Even if it brought back the exact same answers, nothing changed, pay up, right? | 即使它带来了完全相同的答案, 没有什么改变,支付,对不对? |
| 251 | No, it's false positives, for sure go in. | 不,这是假阳性, 当然进去。 |
| 252 | Token costs, right? | 托肯成本,对不对? |
| 253 | We saw in the model that there's a very high token. | 我们在模型中看到 有一个很高的标志。 |
| 254 | We know that doesn't shrink well over time. | 我们知道这不会随着时间而减弱 |
| 255 | There's always some volume element in terms of the cost, but it's not the same as flatlining, | 成本方面总有一些量元素,但与平铺不同, |
| 256 | right? | 对吧? |
| 257 | And if we bring this back to knowledge work, this is where we see teams that are starting | 如果我们把这个带回 知识工作, 这就是我们看到的团队开始 |
| 258 | to cut corners. | 切开角落。 |
| 259 | So I won't research this company every day, I'll look at it once a week or once a month. | 所以我不会每天研究这个公司,我会每周或每月看一次. |
| 260 | I won't ask that question now. | 我现在不会问这个问题 |
| 261 | I don't want all the results, I'll only take 10 results, 20 results, something. | 我不想得到所有的结果,我只拿10个结果,20个结果,什么的. |
| 262 | So we already have the setup. | 所以我们已经安排好了 |
| 263 | We have what we need to do the knowledge work, but at the same time, we're not extracting | 我们有需要做的知识工作, 但与此同时,我们没有提取 |
| 264 | all of the value because we're starting to be conscious about cost, right? | 所有的价值 因为我们开始意识到成本,对不对? |
| 265 | Basically we're renting context. | 基本上,我们租了背景。 |
| 266 | We're not owning the context that we use. | 我们不拥有我们使用的背景。 |
| 267 | That is a very important distinction. | 这是一个非常重要的区别。 |
| 268 | Again, if we're good engineers and we ask ourselves what about scale, the second most | 再说一遍,如果我们是优秀的工程师, 我们问自己,什么是规模,第二 |
| 269 | obvious thing that will come to mind now, so how about we build it? | 显而易见的事情,现在会想到, 那么我们如何建造它? |
| 270 | What if we take all of that web data ourselves and stick it in some vector database and try | 如果我们自己把所有的网络数据 粘在某个矢量数据库里 然后尝试 |
| 271 | and see what comes out of it? | 看看里面有什么 |
| 272 | So I asked my engineer to do exactly that. | 所以我要求我的工程师这样做。 |
| 273 | Again, this is a test. | 再说一遍,这是一个考验。 |
| 274 | This is not a benchmark or a full-blown operation. | 这不是一个基准或全面的行动。 |
| 275 | This is a day's work at best just to illustrate the concept and to show something about the | 这是一个一天的工作,充其量只是 说明这个概念 并展示一些关于 |
| 276 | cost efficiencies that you can generate potentially by doing it yourself, potentially, in specific | 成本效率,你可以 潜在的通过自己做到这一点, 潜在的,具体 |
| 277 | scenarios. | 假设 |
| 278 | The test, simple. | 测试,很简单。 |
| 279 | Take the company name, nothing but run it through Google, find the relevant entries, | 取公司名称,只是通过谷歌运行, 找到相关的条目, |
| 280 | the relevant URLs of that company in various websites that have all of that information. | 拥有所有信息的各种网站上的公司的相关URL. |
| 281 | You use search when you don't know the source, but when we're talking about company enrichment, | 你用搜索 当你不知道来源, 但当我们谈论公司浓缩, |
| 282 | we all know these sources. | 我们都知道这些来源。 |
| 283 | We all know where that data comes from. | 我们都知道这些数据来自何处。 |
| 284 | The cost, also bring it from them, zoom in for bringing it from them. | 成本,也从他们带来, 放大从他们带来。 |
| 285 | It's the same thing over and over again. | 亦复如是. 师曰. |
| 286 | Why not just go straight to the source? | 为什么不直接去源头? |
| 287 | Why are we doing that middleman thing? | 我们为什么要做中间人的事? |
| 288 | LinkedIn公司,LinkedIn 工作,Crunchbase, 在那里,我们有刮刮机的人 | |
| 289 | just tap in and you start paying as a pay-as-you-go. | 随便你便开始付钱 |
| 290 | We built two dedicated scrapers. | 我们造了两台专用刮刀 |
| 291 | 我们有一个新的AI工具,叫Scraper Studio. | |
| 292 | It basically lets you build a scraper for any website in less than five minutes, all powered | 它基本上让你为任何网站 建造一个刮刀 在不到5分钟,所有电源 |
| 293 | by AI, and then it also has a self-healing function. | 由AI进行,然后它也具有自愈功能. |
| 294 | If the website changes, it fixes itself and keeps on going. | 如果网站有变化,它会自我修复,并继续运行. |
| 295 | Merge it all into one entity, basic heuristics. | 把它们合并成一个实体 基本热力学 |
| 296 | If there's conflict, choose that over that. | 如果有冲突, 选择这一点。 |
| 297 | Eventually, we have a data set of these 100 companies, zero AI cost involved. | 最终,我们有这100家公司的数据集,零AI成本。 |
| 298 | There's no tokens. | 无有征兆. |
| 299 | Coverage, fairly well. | 覆盖,相当好。 |
| 300 | Not amazing, not the best that we saw here, but stacking up pretty well. | 并不令人惊奇,也不是我们所看到的最好的, 但堆积得很好。 |
| 301 | Again, this is just a day's experiment, probably not even as much. | 再说一遍,这只是一天的实验, 可能甚至没有那么多。 |
| 302 | Again, very specific tasks, very limited context, very limited situation. | 同样,非常具体的任务、非常有限的背景、非常有限的情况。 |
| 303 | Tread lightly and then proceed with caution when it comes to conclusions. | 轻轻地绊倒,然后在得出结论时谨慎行事. |
| 304 | The real story is not this. | 事实并非如此。 |
| 305 | The real story is this. | 真实的故事就是这个。 |
| 306 | That's what it costs to just go and fetch that data that is out there. | 这就是去获取外面的数据的代价。 |
| 307 | We think about knowledge graphs. | 我们考虑的是知识图表。 |
| 308 | We think about entities, but if you think about, for example, LinkedIn, | 我们想的是实体,但是如果你想, 例如,LinkedIn, |
| 309 | the data is already structured in form of entities. | 数据已按实体形式排列。 |
| 310 | There's an entity for a company, there's an entity for a person, | 公司有实体,人有实体, |
| 311 | there's an entity for a job, and they're connected between them. | 有个实体来工作 他们之间有联系 |
| 312 | Sometimes the ontology is already there. | 有时本体论已经存在. |
| 313 | Again, this is not the most complicated of scenarios, but this is pretty damn good. | 再说一遍,这不是最复杂的情景, 但这是相当他妈的好。 |
| 314 | Now, yes, it took time to set up. | 现在,是的,它需要时间 设置。 |
| 315 | So it's not really fair to compare apples to apples when it comes to the cost, | 所以,如果把苹果和苹果比起来 成本是不公平的, |
| 316 | because these are out of the box. | 因为这些都出柜了 |
| 317 | You can just tap into the API. | 你可以进入API。 |
| 318 | That one that I just showed you required some setup. | 我刚刚给你看的那张 需要一些设置。 |
| 319 | Let's say it's a week. | 如是说礼拜. |
| 320 | Let's price it at $5,000 just to give us some perspective. | 让我们用5000美元的价格来给出一些视角。 |
| 321 | We can actually start thinking about this in terms of a tipping point. | 我们可以从一个临界点开始思考这个问题。 |
| 322 | We can actually start thinking about what is that tipping point | 我们可以开始思考什么是临界点 |
| 323 | in which it makes more sense for me to build it myself than keep on renting it. | 我个人建造它比继续租更合理 |
| 324 | Now again, everything to the left of that dot, in this case, | 再说一遍,这个点左边的一切,在这种情况下, |
| 325 | it was just over 15,000 entities or queries when we think about it. | 我们考虑的时候,只有15,000多个实体或查询。 |
| 326 | So it made sense to do it at this point. | 因此,现在这样做是有道理的。 |
| 327 | Maybe it's not 15. | 也许不是15岁 |
| 328 | Maybe it's 30. | 也许是30岁 |
| 329 | Maybe it's 100,000. | 或谓十万. |
| 330 | Maybe it's 10,000. | 也许是一万块 |
| 331 | It really depends on the use case. | 这真的取决于用例。 |
| 332 | But there is a tipping point in which it actually makes sense to do it yourself, | 但有一个临界点, 它实际上有道理自己做, |
| 333 | which leads us to the fact that both AI Search and CAS and all these solutions, | 这导致我们发现 AI搜索和CAS 以及所有这些解决方案, |
| 334 | they're very good in the sense that you can just plug and play. | 他们很好,因为你可以 插和演奏。 |
| 335 | But if your knowledge work needs are persistent and consistent, | 但如果你的知识工作需要 持续和一致, |
| 336 | and to a certain degree may even continue escalating and growing, | 并在某种程度上可能继续升级和增长, |
| 337 | then this is perhaps a direction to start considering. | 那么这也许是开始考虑的方向。 |
| 338 | Maybe I can just go ahead and build my own. | 也许我可以自己建 |
| 339 | Because the nice thing about it is that all of the things that we see on the left | 因为我们左边看到的一切 |
| 340 | up until the third part is upfront investment. | 直到第三部分是前期投资。 |
| 341 | And the most important thing that whatever retrieval happens later on from the agents | 最重要的事情是,不管后来从特工那里收回什么 |
| 342 | is free. | 是免费的。 |
| 343 | Not really free, but you get what I mean. | 不是免费的 但你懂我的意思 |
| 344 | There's no added cost. | 没有额外的成本。 |
| 345 | I can just ask that question over and over again. | 我可以反复地问这个问题。 |
| 346 | I did not like the first answer. | 我不喜欢第一个答案。 |
| 347 | I'll ask it again. | 又问. |
| 348 | I'll ask it 100 times until I get what I need. | 我会问100次 直到我得到我需要的。 |
| 349 | I have no more fear, no more cutting corners, which is maybe the most important thing. | 我不再害怕 不再割角 这也许是最重要的 |
| 350 | And I'm leaving aside the fact that this is also custom business logic. | 我撇开这个事实, 这也是定制商业逻辑。 |
| 351 | I can connect it with my own data. | 我可以用我自己的数据连接它 |
| 352 | There's all sorts of other advantages of owning it. | 拥有它还有其他各种好处. |
| 353 | We'll keep it to the imagination. | 我们会保持它的想象力。 |
| 354 | Remember, we asked about a million, not about 15,000. | 记住,我们问了一百万,不是15,000。 |
| 355 | This compounds. | 这种化合物。 |
| 356 | This compounds greatly. | 这非常复杂。 |
| 357 | We need that horizon. | 我们需要那个视野。 |
| 358 | Remember, the web keeps changing. | 记住,网络一直在变 |
| 359 | We saw the staleness of the data and how the data decays. | 我们看到数据的停滞 和数据如何衰减。 |
| 360 | So we need to be thinking about this in the long run and how this will evolve | 所以我们需要从长远的角度来考虑 这个问题会如何演变 |
| 361 | when we keep on asking the questions about the entities that we care about. | 当我们不断问我们关心的实体的问题时 |
| 362 | Just to wrap it up. | 只是把它包起来。 |
| 363 | So AI, Search, CAS, they can get you very far when what you need is ad hoc | 所以,AI,搜索,CAS, 他们可以让你非常远 当你需要什么 临时 |
| 364 | and what you need is always changing. | 你需要的总是在改变 |
| 365 | When sometimes you look at different things, | 有时你看着不同的东西, |
| 366 | even the mix and match of them for certain tasks use this, | 即使混合和匹配 某些任务使用这个, |
| 367 | for certain tasks use that. | 用于某些任务。 |
| 368 | You can, I'm sure, again, I just saw the test. | 你可以,我敢肯定,再次, 我刚看到测试。 |
| 369 | There's a lot of ways to optimize it just like any other context engineering | 有很多方法可以优化它 就像其他环境工程一样 |
| 370 | and use lighter models and use other stuff. | 并使用更轻的模型 并使用其他的东西。 |
| 371 | There's a lot of great stuff to be done. | 众生种种妙法. |
| 372 | But eventually, the frequency will come and bite you in the ass | 但最终,频率会来咬你的屁股 |
| 373 | when it comes to cost. | 当涉及到成本。 |
| 374 | And that's something to be mindful of. | 这是值得注意的。 |
| 375 | And there's a fair chance that that tipping point is much lower than you think. | 并且有相当的机会, 这个临界点 比你想的要低得多。 |
| 376 | And that's something that, as we design these systems, | 当我们设计这些系统时, |
| 377 | when we think about web context engineering, we need to be mindful of that. | 当我们考虑网络环境工程时,我们需要注意这一点。 |
| 378 | And last but not least, the last slide we showed. | 最后但并非最不重要的是,我们展示的最后一张幻灯片。 |
| 379 | Owned context compounds while rented decays. | 租赁衰变时拥有上下文化合物。 |
| 380 | It's not a one-time task. | 非一时之任. |
| 381 | Again, if it's a one-time question, use AI, Search, it will be amazing. | 复次若是一时问,用AI,Search,会令人惊叹. |
| 382 | When you need to do it over and over again, | 当你需要一次又一次地做的时候 |
| 383 | there's a fair chance that it will not, | 很有可能不会 |
| 384 | that you're missing out on potential compounding effect | 你错过了潜在的复合效应 |
| 385 | and you are losing out. | 你输了 |
| 386 | Thank you very much. | 谢谢 |
该视频共有字幕 386 条。解锁更多字幕为会员功能,请移动到 价格
