Latka logo

Featherless AI Revenue & Funding (2026)

Featherless AI is the largest open-source LLM inference provider on Hugging Face, founded in 2024 and led by CEO and co-founder Eugene Cheah. The company offers serverless access to over 30,000 open-source models on a flat-rate pricing model, abstracting GPU infrastructure complexity for developers and prosumers who want to run AI models without managing hardware. Cheah co-created RWKV, the first attention-free AI architecture under the Linux Foundation, and the company grew out of that research community as an accidental pricing experiment that quickly outpaced the original platform.

As of April 2026, Featherless AI serves 10,000 customers, generates more than $250,000 per month in revenue but less than $500,000 per month, and has a team of 27 people approaching 30. The company raised a $2 million seed round pre-launch for RWKV research, then closed a $20 million Series A in December 2025 led by Airbus Ventures and AMD Ventures. Its largest single customer pays between $1 million and $2 million per year.

The single most important strategic fact is that Featherless AI built its inference stack from scratch to load models in 5 to 30 seconds, versus the industry-standard 30 minutes, enabling dynamic GPU swapping across thousands of models on demand. This technical moat lets the company serve the long tail of fine-tuned and niche-language models that no other provider hosts, which is what drove the initial Reddit-fueled signup surge and continues to differentiate it from competitors who cover only the top 100 models.

Last updated

Featherless AI Revenue

As of April 2026, Featherless AI generates more than $250,000 per month in revenue but less than $500,000 per month, with the company targeting $500,000 per month by fall 2026. Scale-up customers, those requiring 100 or more concurrent requests, represent roughly half of total revenue, while the entry-level $25 per month tier drives the individual user base.

Featherless AI Revenue GrowthReported revenue / ARR over time$0$100K$200K$300K$400K$500K202420252026$0$400K$250KSource: GetLatka.com interview on Apr 24, 2026 with Featherless AI CEO Eugene Cheah
YearMilestoneSource
2026Featherless AI Hit $250k revenue in April 20269:42[1]
2025Featherless AI Hit $400k revenue in April 2025
2024Launched with $0 revenue

The company grew from zero to its current revenue run rate within roughly two years of launching in 2024, with early growth so rapid that server capacity was continuously overwhelmed until the December 2025 Series A provided enough capital to stabilize infrastructure. The largest single customer pays between $1 million and $2 million per year, and the company is actively negotiating additional multi-million-dollar annual enterprise contracts.

Featherless AI Valuation, Funding Rounds

Featherless AI has not publicly disclosed its valuation. The company has raised $2M in total funding to date.

Featherless AI has raised $2M in total funding across 1 round, most recently a $2M Seed round in 2023.

Featherless AI Capital Raised & ValuationCumulative capital raised and post-money valuation by roundCapital raised (cum.)Valuation$0$0$0.2$500K$0.4$1M$0.6$1.5M$0.8$2M$1$2.5M20232024Source: GetLatka.com interview on Apr 24, 2026 with Featherless AI CEO Eugene Cheah
YearRoundAmountValuation% SoldSource
2023Seed Round$2M--

Founder / CEO

Eugene Cheah

CEO

Eugene Cheah is the CEO and co-founder of Featherless AI. He is one of three co-founders, alongside Wesley George and Harrison Vandevelle. Cheah co-created RWKV, described as the first attention-free AI architecture, developed under the Linux Foundation. RWKV has the potential to reduce inference costs by over 1,000 times compared to current architectures, according to Cheah, though he noted the architecture has not yet been proven and scaled to the largest model sizes.

Prior to Featherless AI, Cheah operated under the entity Recursor, which ran an inference platform for RWKV models and conducted foundational model research. Recursor was exited in 2024. Featherless AI began as a pricing experiment under a separate brand name while Recursor was still operating, and within its first few days generated more revenue than the original Recursor platform, prompting the team to pivot fully to Featherless AI.

Cheah's stated motivation for building the platform is rooted in access: he cited his grandmother, who speaks Thai, Cantonese, Hokkien, and Bahasa Malay but not English or Chinese, as an example of the billions of people who would be excluded from AI if only mainstream English-language models were supported. Net worth was not discussed in the interview.

Q&A

QuestionAnswer
What's your age?-
Favorite online tool?-
Favorite book?-
Favorite CEO?-
Advice for 20 year old self-

Customers

Featherless AI serves 10,000 customers as of April 2026, ranging from individual developers and prosumers on the $25 per month entry plan to enterprises paying between $1 million and $2 million per year. The customer base is split between self-serve users who discover the platform through Reddit and Hugging Face and high-touch enterprise accounts closed by the sales team.

The company's go-to-market positioning targets companies spending $100,000 or more per month on closed-source API credits from OpenAI or Anthropic, offering them open-source alternatives that can significantly reduce inference costs. Featherless is listed as a filterable inference provider on Hugging Face model pages, giving it organic discovery placement for the niche and fine-tuned models where it is often the only available provider.

Featherless AI serves 10K customers.

Featherless AI Business Model

Featherless AI operates on a flat-rate monthly subscription model starting at $25 per month, with usage-based scale-up plans for customers requiring higher concurrent capacity. The company hosts only open-source models, which eliminates per-token licensing costs from closed-source providers and allows the flat-rate structure to remain margin-positive. GPU costs cited by Cheah range from $5 to $80 per hour depending on server size, with a 200-billion-parameter model such as Step-1 requiring at least four H100 GPUs at approximately $20 per hour.

The company's core technical advantage is its proprietary inference stack, which loads models in 5 to 10 seconds for standard models and under 30 seconds for especially large models, compared to approximately 30 minutes for traditional inference infrastructure providers. This fast load time enables dynamic model swapping across a shared GPU pool, allowing Featherless AI to support 30,000 models currently and target 100,000 without maintaining a dedicated GPU for each model simultaneously. Cheah stated that the top 100 models account for 50 percent of the company's inference workload, with the remaining 50 percent spread across long-tail fine-tuned and specialized models that most competitors do not support.

Featherless AI claims its pricing is at least 10 times cheaper than closed-source inference providers. The RWKV architecture underlying some of the company's optimization work has the theoretical potential to reduce inference costs by over 1,000 times, though Cheah noted this has not yet been demonstrated at full scale. The company rents the majority of its GPU capacity from third-party data centers in North America and Europe rather than owning hardware. Profitability was not discussed in the interview. Gross margin, churn, LTV, CAC, and burn rate were not discussed in the interview.

Point-in-time figures shared on the GetLatka podcast, each linked to the exact moment it was said on camera.

Customers (2026)

10,000

Eugene Cheah: Yeah, so we are serving around 10,000 plus customers and I would jokingly say the average customers do not know what a B200, MI325, they may have heard of H100 because it was in the news.

Watch at 8:46

Featherless AI Employees & Team Size

Featherless AI had 27 full-time employees as of April 2026 and was actively hiring toward 30, which matches the current team size reported in company data. The team is divided into three functional groups: approximately 12 people on infrastructure, 10 on platform and go-to-market including developer relations, and the remainder on research and marketing activities.

The research team continues active work on the RWKV attention-free architecture in parallel with the commercial inference platform, reflecting the company's dual identity as both a research organization and a production infrastructure provider.

Featherless AI employs approximately 30 people as of 2026, up from 17 in 2025. It serves 10K customers that rely on its solutions.

Featherless AI Team GrowthReported headcount over time08152330382024202520260017173030Source: GetLatka.com interview on Apr 24, 2026 with Featherless AI CEO Eugene Cheah
YearMilestoneSource
2026Reached 30 employees (April 2026)
2025Reached 17 employees (July 2025)

Frequently Asked Questions about Featherless AI

What is Featherless AI's revenue?

Featherless AI generates $250K in revenue.

Who is the CEO of Featherless AI?

The CEO of Featherless AI is Eugene Cheah.

How much funding does Featherless AI have?

Featherless AI raised $2M across 1 round.

How many employees does Featherless AI have?

Featherless AI has 30 employees.

Where is Featherless AI headquarters?

Featherless AI is headquartered in San Francisco, California, United States.

Full Interview Transcripts

Apr 24, 2026

Nathan Latka (00:01) Hey folks, my guest today is Eugene Chia. He's the CEO and co-founder of featherless AI, the largest open source LLM inference provider on hugging face. He's offering serverless access to over 6,700 models on a flat rate pricing model that cuts inference costs by at least 10 X. He co-created RWKV, the first attention free AI architecture under the Linux foundation. Eugene, you ready to take us to the top? Eugene Cheah (00:26) Yeah, when you look at this fundamental piece of technology called AI, it's something that we believe that shouldn't be controlled by only a handful of companies where they can choose to restrict your access and what you do with AI. So because of that, we fundamentally believe that people should be able to make their own choices and decisions when accessing AI models. And the best way to do it is to support all of them in the open source ecosystem. And that's why we built Federalist AI to support any AI model that you have on hugging face and to provide instant access to all of that. Nathan Latka (01:04) And so here's a list of a lot of those models for the non-technical listener that's using right now. Can you dumb this down for a second, explain it like you're explaining it to a kindergartener. Eugene Cheah (01:15) So AI models can be used for various use cases. You have the typical chat GPT use case. You can have the AI use cases for supporting people in particular language or domains. So ⁓ for example, are AI models specifically tuned for the use of providing agriculture advice for farmers in both the the Asia region and also a few models specifically for the North American regions. There are also AI models specifically Nathan Latka (01:46) Is this an example, Eugene? I'm gonna give a visual here as you describe it. So if I type agriculture here and hugging face, these some of the examples you're talking about? Eugene Cheah (01:56) Not exactly, but yeah, so for example, if you search the one, sorry, I will need to find the examples all over the head, but you can actually search more specifically like languages. So if you look on the left hand side on hugging face for languages, you can actually filter AI models by language itself as well. like something is easy for us to forget that. that the English-speaking nation is less than half of the world population. But there is actually a lot more people out there. And that's what kind of like got me to work on supporting AI models in languages and use cases beyond the mainstream. Like for example, my grandma, she speaks seven languages. She lives in Malaysia, but English nor Chinese is neither one of them. And if you follow some of the mainstream models, she wouldn't be able to speak to AI. And that's what I fundamentally want to prevent from happening, to ensure everyone in the world have access to AI. Nathan Latka (03:02) What language does your mother speak? Eugene Cheah (03:05) So my grandma, she speaks Thai, ⁓ Cantonese, Hokkien, and a little bit of Bahasa Malay, ⁓ respectively, because she's in Malaysia, and those are her main languages. Nathan Latka (03:21) So let's use tie. Let's use tie for example here. So I'm in languages. I just went to tie and I see like there's one, there's, one called text classification, 67,900 downloads. Is this an example of sort of what you mean? Eugene Cheah (03:34) Yes, so but this is a text classification model. You may want to actually use more of a text model itself. ⁓ So for example, one of popular models in Asia is the sea lion model, ⁓ or SEA. SEA, yeah. So ⁓ you may need to remove the language filters, the reset filters, yeah. So this is for example one of the lines of model that supports most of the major Southeast Asian languages. Nathan Latka (04:14) So give me an example of how your mom would use this or how this enables someone to build a model or a version of ChatGPT maybe that your mom could use. Eugene Cheah (04:24) So they could use existing applications or even go to Feather's AI inbuilt application as well. We have a chat application at the top there. And they could actually assess one of the many models ⁓ that we have in our catalog. And this includes some of models that we have highlighted respectively. Usually what most people will, what we see from the community is that like be it through the redates or local communities that people will actually find their own preferred model. And they'll know the models that they would like to run and they will actually come to our platform and they'll make the request and we'll add support for it and then they can run it straight away. Nathan Latka (05:05) So let's use the, it's sorted right now downloads high to low. Is this pronounced quen? Quen ranks number one in terms of number of downloads. Eugene Cheah (05:13) Yeah, this is one of the most popular downloaded models. So when the user signs up on the platform, they can have access to any of these models. So Quen in particular is ⁓ one of the latest ⁓ models made by Alibaba. And it's ⁓ strong in both English and Chinese and a few other ⁓ traditional languages. Nathan Latka (05:34) You guys can read the description of sort of what it does up here. So let's keep going down the stack here, Eugene. My goal on this interview is to help take a very technical sort of concept and help my listeners understand why will you build up featherless.ai is so important in terms of the end value, right? So are mostly, are, it mostly developers that are paying you for featherless? Eugene Cheah (05:53) Yeah, so developers, we also see an increased wave of what I call prosumers, and to us who's like not exactly purely developers. Some of them would be like coming in and it's like, I want to run my cloud code or I'm white coding my apps and I want my apps run on AI. So these are the two major categories. Traditional developers are building apps that uses AI and these groups are and to us as well. Yeah. Nathan Latka (06:19) And you have multiple different price points here, but what would you say the average customer is paying you per month today? Eugene Cheah (06:25) So the average entry level customer is paying the $25 a month plan. That provides them access to any of the model where they can make unlimited requests, limited to one request at a time. And then subsequently, once they figure out which models they want, and sometimes when they create apps and ship it to the app store, that's where they go to our scale up plan. And that's where they have much larger dedicated capacity because you don't really want your production apps to Nathan Latka (06:50) Who's paying for the credits though there? If they pay you 25 a month and then they use some of the models that you help them sort of play with and then it gets used a bunch, who's paying for their credit? Surely they're paying for their credits. They're putting in their API key, I guess, and they're paying for the credits? Eugene Cheah (07:03) Yeah, so we provide the entire infrastructure and the infant service. So the $25 a month is paid directly to us. And then subsequently, they are able to access the growing catalog of 30,000 models, which we are scaling respectively. Nathan Latka (07:18) How do you manage the margins if they're using a model that charges you like a dollar per API call but they're only paying you $25 a month? Don't you go underwater on their plan? You're losing money on their plan once they do 25 API calls? Eugene Cheah (07:30) So to be clear, we don't serve OpenAI or Anthropics closed source models. We are hosting the catalog of open models. Nathan Latka (07:40) only open source. So they don't charge per credit, these open source models. Eugene Cheah (07:41) Correct. ⁓ Correct. And these are models that you could have downloaded on your own high-end GPU if you have the GPU and run it. But in our case, we will just load it onto our servers when you make that request. So essentially, that $25 a month is you're getting a fraction of the GPU compute ⁓ that you have access to, and then we'll run the models on it. Nathan Latka (07:47) Got it. This makes sense. This is like back in the days when everyone went out and bought server racks and put it in their closet, right? Versus now just using AWS. You're sort of the modern day version of that for AI open source models. Eugene Cheah (08:22) Correct. So for example, like how Heroku or Versailles abstract the whole infrastructure layer for just deploying applications, we are abstracting that for AI models. And if you want to host and run at larger scale, we also provide the pricing plans for it. Nathan Latka (08:38) I see, very cool. Okay, it's now making sense to me and I'm not a tech person, so I imagine my audience is now following along nicely. How many customers are you serving now today? And then we'll get your backstory here. Eugene Cheah (08:49) Yeah, so we are serving around 10,000 plus customers and I would jokingly say the average customers do not know what a B200, MI325, they may have heard of H100 because it was in the news. And that's essentially what our job is. We abstract away all the complexity of running these AI models ⁓ and the infrastructure for it. ⁓ and these growing collection of open models include some of the best models that are already on par or surpassed with let's say Cloud Sonnet or even JPT for our mini. And we realized actually for a lot of customers when they move to production, these are the models that are more than sufficient for their needs. Nathan Latka (09:35) Interesting. And just to be clear, if you've got 10,000 customers paying that average monthly price point, you just told me of 25 a month. Eugene, can I multiply those together to get your monthly, you know, your monthly revenue is about 250,000 a month. Is that accurate? Eugene Cheah (09:49) No, it's higher than that because those will be for more towards individual users. But for the scaled up customers and for the scaled up customers, which actually represents more like half of our revenue, these customers will be coming in and saying, hey, I would like to have 100 concurrent requests. I would like to have 1,000 concurrent requests, which is a much higher volume than these individual users. And this is a pattern that we see increasingly where where developers will come in, test, prototype, find something in like, and scale up. Nathan Latka (10:23) Eugene, I'm super impressed. So sorry, just to be clear, I don't want to put you on the spot. I didn't prepare you for this, but are you comfortable sharing your monthly revenue today? Is it more like $500,000 a month? Eugene Cheah (10:33) ⁓ That's something that we are scaling up towards ⁓ and we are not yet at that number. these are those things that we already like right now the inference market has been growing rapidly that some enterprise customers that we are talking to, we are negotiating like multi-million dollar contracts on an annual basis. So it's something that we can probably hit within the next few months. Nathan Latka (10:58) That's exciting. just to be clear, we're recording in April of 2026. You're doing more than $250,000 a month, but less than $500,000 a month, but you're growing so quickly. You think you'll pass $500,000 a month sometime in fall here of 2026. Eugene Cheah (11:14) Yeah, AI is really exciting, especially in this space where some companies literally are renting entire clusters of AI servers to automate their workflows. just renting even two of these biggest servers at full capacity, let's say the B300 or the MI325 for the AMD side, are already with a million dollar contract per year. Nathan Latka (11:42) Are you holding this inventory on your balance sheet? I mean, are you running a physical plant holding all these things somewhere or are you renting these from somebody else? Eugene Cheah (11:54) So for the vast majority of our inventory, we do rent it from someone else. ⁓ So we rent it directly from the data center. But once again, I say, our average customers, do not know how to set this up or configure or manage it. And for the data center providers, they don't really want to take the responsibility of standing up the entire software infrastructure stack. And that's where, for some of our private cloud or dedicated customers, We handle all of that for them. The public cloud that you see is just one way for the customers to try and experiment. But when they scale it up, we can either provide from our inventory or even use the customer's inventory. Nathan Latka (12:37) Eugene, tell me more of your backstory here. We need to write the first line of code for Featherless. What year? Eugene Cheah (12:43) So, Fedulus started out kind of like as an accident that outgrew itself two years ago. So, as you heard, ⁓ we started from the RWKB committee where we were doing experiments in next generation foundation models. And that's where our roots are. Like this new AI architecture has the potential of reducing inference cost by over a thousand X. And if you hear all the energy demands, that is extremely lucrative. The downside is that this new generation of AI architecture hasn't been proven and scaled to the largest model size. so, correct. So multiple members of the team was working on this project. And so during the early stages of the inference platform that we did, well, this shows a lot of potential and we are still scaling it to eventually replace all the AI models you see today with something much cheaper. We created a platform where Nathan Latka (13:20) That's this, just to be clear, right? Eugene Cheah (13:41) you could fine tune our models for specific tasks. And this is what people use small models for. Because if you use small model, you have to fine tune for a specific use case. We end up having so many people fine tuning on our platform that we needed to create a platform, inference platform that can support thousands of models. And then when we end up creating that, we're like, why are we only running RWKB here? So we added support for all the other models. And then suddenly, All the signups happened and we are like, we used to say on our website, featherless, Instant access to thousands of models ⁓ made cheaper and powered by RWKV because there's some truth to it. Like underneath the hood, we use RWKV to optimize the inference, but we actually realized it lowered the conversion rate. We realized people just want, I want this model, I saw your price tag, that's all I need to know. Nathan Latka (14:32) That's amazing. So you can see our RWKV here inside of now Featherless. But what you're saying is you effectively were doing all the research here. You built this Featherless for yourself and you said, my gosh, everybody else needs this as well. What was the first massive sort of signup surge that you saw? Was it a article on, on hacker news or somewhere else? What got, or Reddit, what got you the first bump of signups? Eugene Cheah (14:56) So when we did the first soft launch, we actually just, we had a few members of the committee just posted on Reddit essentially, to some of the existing AI committees, ⁓ because we knew that we wanted to serve the models that no one else hosted. And that's what really drove the traffic. ⁓ You see most providers, they only provide, let's say, less than 100 models. That covers 50 % of our inference workload. It's the bottom 50 % where they run all these interesting fine-tuned models that people came on board for. And they are all usually very unique use cases or languages. Nathan Latka (15:37) Is this you? Is this your silly tavern AI? that you? Eugene Cheah (15:41) That is one of our members that probably did the initial post. We have three founders. ⁓ So this is probably it was done by Wes. So Wes, Wesley George was our CEO. He's based in Toronto and Harrison Vandevelle. He's based in Australia, our CTO. Nathan Latka (15:47) co-founders. How many co-founders do you have? You Okay, wow, so 2024 was official launched in two years ago. And you're doing more than $250,000 a month today in revenue. So it's fair to say you've gone from zero to a million dollars of revenue. What like very quickly, right? In a couple of months. Eugene Cheah (16:19) Mm, mm, yes. So, so. Nathan Latka (16:22) You're like, you're an engineer and I'm a business guy. So it's uncomfortable when I ask you finance questions, but I love to capture the growth story. Eugene Cheah (16:30) You want to hear the funniest bit about this right? The company at that point of time was called Recursor. So we already had an inference platform for RWBKB. Recursive model, RWBKB. It all makes sense. And Federalist was meant to be a pricing experiment. So we gave it a different name. But within the first few days, it became more profitable and more revenue than the original company platform that we were like... Nathan Latka (16:34) Yes. Okay. Eugene Cheah (16:59) I guess we are featherless now. Nathan Latka (17:03) That's amazing. what did it say? mean, did you guys go, I mean, you guys, as your co-founders, you must have said, oh my gosh, we just passed 88, $83,000 a month in revenue and we've only been live for like three months. How many months did it take you to break 83,000 a month? Do you remember? Eugene Cheah (17:18) wow, I can't really remember that moment but it was like... It was all such ⁓ a blaze because like it was a case of like we get more users, the servers are on fire, we add more servers, we get more users, we add more servers and ⁓ so it was like just a constant hectic rush there. And it wasn't until like ⁓ much more recently where we had a lot... we recently... ⁓ had a close our funding our latest a round where we had where we had it now like enough server capacity is like we have a sign already now the there is slightly more servers and users for now Nathan Latka (18:01) for now. That's great. Tell me tell me more about the funding history. You just said you just closed a series a when did you close that and how much was it for Eugene Cheah (18:10) So we closed our series A ⁓ sometime during December. That was a $20 million round. ⁓ And the leads for it was Airbus Ventures and AMD Ventures. And since then, we've been scaling the platform much more aggressively. Nathan Latka (18:28) So you closed 30 million series a in December of 2025, 20, 20 million. Okay. And what about the seed? Did you do a seed round or did you skip it? Eugene Cheah (18:31) No, 20. Yeah. We did do a seed round much earlier and ⁓ that was pre-Fedellus. That was when we did a seed round over the research of RWKB. Nathan Latka (18:48) and how much was the seed round for? Eugene Cheah (18:51) The seed round was a small 2 million round. Nathan Latka (18:55) Okay, this is great. How much would I have to bribe you to let me put in 100K on the Series A round and I'll promote the hell out of this to drive you guys more users. Do I have to bribe you? Do you take bribes? Eugene Cheah (19:06) Ha We accept users. Nathan Latka (19:11) ⁓ So there's zero chance that I can write a 100k check on the Series A docs. I won't even argue about valuation. I won't even do any diligence. I close my eyes, I write a 100k check, and you and your co-founders take it. Eugene Cheah (19:26) ⁓ I mean we do have some flexibility to do a safe note here and there as part of this round but yeah, as long ⁓ as you can help us get users, think our investors are more than happy to entertain the idea. Nathan Latka (19:41) Well, I'm a marketing guy and we've invested 250 million and 700 portfolio companies that I think should all be looking at what you're building. So I'll take this offline with you to try and convince you to let me put in 100k on a safe note. You can make up whatever valuation you want because I think this is, I don't understand everything you're doing. I'm not a technologist, but I love how you talk about building the product and the growth speaks for itself. Eugene Cheah (20:03) I'm all ears for that and I would like to know how many of your portfolio is burning so much AI usage that we can come in to help them lower their costs. Nathan Latka (20:12) Tons. mean, this is why it's interesting, right? When we look at all of the profit and loss. So every portfolio company has to connect their profit and loss to founder path. I can look in their cost of goods sold line and I see how much they're paying to Anthropic and open AI and on all the credit spend. If you're telling me that you've got a way to help them cut their cogs in half, right? Or even more, that's extremely valuable. Eugene Cheah (20:37) Exactly, and that's actually how our sales team are starting to close a lot of these deals because they'll come in and say, hey, what are you using AI for? Do you know for half of this workload, you could use this open model that's much cheaper and lower? We can provide that. For this half of this work, you can use this model. And since we have seen them all, we can advise more specifically. Nathan Latka (20:59) I love that. how much of what you do, would you say is sort of note people buying you sort of without emailing you without a call versus high touch you telling people what model they should use or how they can save 10 times or, 10 X their inference spend. Eugene Cheah (21:12) It's it's currently a bit of a both. So when it comes to the individual users, when they're coming in on the public cloud, this is true word of mouth self discovery for most of the cases. And every now and then some of these users will upgrade down the path respectively. But we also realize that there is a lot of money on the table right now where you can go after the startups that, hey, I just built my entire startup or SMB. on OpenAI or Entropic and I'm $100,000 a month and ⁓ I do not know what I was doing. And it's like, okay, we can help you here. We can lower your bill by half and we can see where it goes. That is usually some of most ideal large volume customers ⁓ when they are basically spending that much and they are entering a situation where, hey, we need to start thinking about this and then we step in. Nathan Latka (21:50) Yep. Yes. Okay, so here's a deal. I know you're an engineer, you're not a deal guy, but I'm gonna try and sell you on a deal. If I can bring you five customers that are spending more than 100 grand a month paying for credits for the models and they sign up with Federalist, you'll let me put $100,000 in on a safe note. I don't care what the valuation is. Do we have a deal? Eugene Cheah (22:26) I think I sometimes have a hard time. Nathan Latka (22:31) I love this show. is great. Tell me more about your team, Eugene. How many people are full time today? Eugene Cheah (22:40) So we currently have around 27. We are close to past the 30 mark soon. We are aggressively hiring. ⁓ And our team is split across both the platform for deploy engineers, which supports the customers, the infrastructure team that keeps everything running but doesn't do anything with building, for example. And then we also have the research team that is still working on the other RWKV line of models because we still feel that there's a lot of room still to optimize these AI models today. Nathan Latka (23:14) how many of the 27 are working on infrastructure. Eugene Cheah (23:19) ⁓ So ⁓ around 12 is working on infrastructure right now. ⁓ And then 10 is moving towards platform, go to market. And then the ⁓ rest is ⁓ increasingly like GTM, like marketing activities and research. Nathan Latka (23:25) Okay. Okay, got it. So 12 on infra 10 on like platform and then the rest, you know, call it five, six people are sort of research and go to market motion. Okay. Eugene Cheah (23:50) Yes, the platform does do some marketing as well. They do more like like dev rel and things like that. Nathan Latka (23:58) And is all of the inbound right now pure word of mouth or just Reddit or how are the customers finding you? Eugene Cheah (24:06) ⁓ Both via Reddit, Huggingface, various other platforms. We have started trying to, we also have started like preparing marketing materials, but those haven't really kicked off yet. So it's mostly been Reddit and what not. Nathan Latka (24:22) Where do people find, like I'm on hugging face right now. Where could I click on hugging face to see, hugging face promoting featherless. Eugene Cheah (24:28) So if you go to Models, so this is a little bit tricky because it's, and you go under Others, oh there, you can see it appear. There's a pop-up there. You can see Federalist AI. You can filter the inference provider. Federalist AI. Yeah, so you select Federalist AI on the left-hand side. Nathan Latka (24:36) under Where? Filter by inference available. Featherless, why don't I see featherless? Eugene Cheah (24:55) You just went over it. Yeah. Yeah. Yeah. So you reset all your inference provider filter and select only $5. So what happens is when you come to some of these models, so for example, you can look at the... Let's take the recent one. That starts with an nand bridge because Quen is quite popular. A lot of providers are providing it. But if you go down the list slightly, Nathan Latka (24:56) here, here. Eugene Cheah (25:23) One of the step fun models, for example. Nathan Latka (25:28) You're saying this this step the what the step what? Eugene Cheah (25:31) Step fun, yeah that model. For example, if you click on that, ⁓ we are one of the only providers online right now for this model. Nathan Latka (25:33) right here. ⁓ so this is the pop-up here that you're getting promoted. Eugene Cheah (25:47) Yeah. Yeah. So if you come, if you, if, if you are paying card game fees on the pro accounts, you can actually instantly do inference on fair dollars for these models. So I'll, I'll go to market strategy competition here, right? If you're not going to fight the giant battle of like the top 10 models. Quinn is one of the the top 10 and top 100 models where all the providers are fighting it out there. You're going to see like eight, 10 providers. Our strategy is Nathan Latka (25:49) I see. Eugene Cheah (26:16) We are supporting all the other models that people are interested in experimenting. And for example, the step-1 model is a particularly popular model for us that easily shipped several contracts for us on this model alone. And no one else is doing it. Nathan Latka (26:33) Why are you guys, I mean, I know you're smart, right? I I don't understand all the research, but I can see you're doing a ton of research, but why are you the only infants provider listed here? Is it just really hard to build that inference model? Eugene Cheah (26:47) Yeah, so right now today traditionally when you do inference as an infrastructure provider, when you load a model onto GPU and set it up, it takes around 30 minutes, give and take, for the models to load and get up and running. That 30 minutes is cost time. When we're talking about GPUs that cost, let's say, even $5 an hour or an entire server that can cost $80 an hour. You don't want to half an hour here, half an hour there, letting it run just to load a model. So operationally, as an infrastructure company, you do not, an inference company, you do not want to be switching out models left and right to cater to various user demands, because you're going to waste all this switching time. We built our inference stack and our infrastructure design from scratch where we load models in. 5 to 10 seconds or 30 seconds and under for the especially large one. And what this lets us do means we can actually change our GPU to serve the various demands for various models live. And that allows us to host a lot more models. When we say we support 100,000 models, right now 30,000, it doesn't mean we have 100,000 GPUs all the models live right now. It just means we have, let's say, a few thousand GPUs. hopefully a lot more as we have a lot more customers. And when they ask for a model, we load it into the GPU on demand. Nathan Latka (28:24) And that on demand is what makes you unique and the only reason you can do that is because you built your whole stack from scratch. Eugene Cheah (28:32) Correct. We heavily customized our stack from the get-go to specialize into being able to swap and load these models dynamically. Nathan Latka (28:41) So is it fair to say then if I go back, mean are these your biggest competitors over here? Eugene Cheah (28:50) Yeah, when you talk about the top 10 or top 100 models, yes they are. If you talk about the rest, that's where we come in. And in an AI landscape where, try to view it the other way. Like today, a lot of companies have started fine tuning their own models for their own unique use cases and specialization. ⁓ In a world where various companies are fine tuning their own model. You can't go to a provider that can only support 100 models. There are more than 100 companies on earth. You want an infrastructure tailored to be able to handle all these various fine-tunes. Nathan Latka (29:30) So do you have the largest coverage? mean, is that what you measure? How many models can you cover? Eugene Cheah (29:36) Yes and that also allows us to have a lot of demand for all these models like like step fun for example is not an unpopular model it's shipping billions of tokens per day. Nathan Latka (29:50) How do you know that? Can I see that somewhere on Hugging Face? Eugene Cheah (29:55) Yeah, unfortunately, I don't think Hanging Face provides that statistics, but if you go by the download count, it's quite a popular model. You have to understand that this is a 200 billion parameter model, meaning you need at least some of highest end GPU. We're talking about at least four H100. We're talking about $20 per hour systems to run this model. And we're talking about hundreds of thousands of companies are already using this model itself. Nathan Latka (30:25) And give me an example just so my audience understands. mean, can you, like, do you know the physical location of the data center that you're renting space from? Eugene Cheah (30:33) Yes, ⁓ so we primarily ran from ⁓ North American data centers where we from a few key partners, we are increasingly also purchasing capacity in Europe from a few key partners and providers ⁓ respectively ⁓ because ⁓ for the sovereign AI demand in Europe, we are actually getting more and more customers that say, I only want my AI to be within a certain ⁓ sovereignty and we are actually helping facilitate that as well. Those are our key major regions. Nathan Latka (31:05) And so it sounds like your biggest monthly expense, maybe you're spending north of a hundred thousand dollars a month on renting these different data center spaces. Is that, that accurate? Is that one of the reasons you needed to raise capital so you could go buy space faster? Eugene Cheah (31:21) Yes, and also one of the things is to actually to get direct access to and from the suppliers as well because I think that's one of the most exciting thing about, would call it a happy accident that where we got investment from AMD was that if you kept track of the GPU space or even compute space in general, in the past six months, we went straight to inventory being wiped out. Not just on the GPU side, on the CPU side, you may have heard within the consumers, people are complaining about RAM prices shooting up through the roof. Supply has been heavily eaten up, but by raising capital from strategic investors like AMD in particular, and a few other who manufacture some of the GPUs and the systems as well, and ⁓ Airbus Ventures and a few other Nathan Latka (32:10) Like this one here, right? AMD. Eugene Cheah (32:21) ⁓ investors in the industry, we are able to get access to the supply without going through all these hurdles right now. So that has been growing advantage for us as well. Nathan Latka (32:34) Interesting, got it. Yeah, so you've picked investors that can also help you manage your cost and accessibility to this sort of data center processing that you need. Yeah, interesting. Very, very cool. Okay, I've learned a lot on this episode. I guess let me just ask one or two other questions. I know Coheer just from my background in the SaaS space. I don't think there are technologists like you. Why do they have an inference provider option over here? Eugene Cheah (32:45) Correct. Exactly. So Cohere in particular, ⁓ they have their own particular line of models that they created ⁓ for basically the North American market and in particular Canada. So they are going to the direction of highly tailored sovereign AI models ⁓ for the domestic market. And we actually see this happening more and more. So for Cohere, they will service the Canadian market. The US market is going to be served by...

Featherless AIApr 15, 2026

Nathan Latka (00:01) Hey folks, my guest today is Eugene Chia. He's the CEO and co-founder of featherless AI, the largest open source LLM inference provider on hugging face. He's offering serverless access to over 6,700 models on a flat rate pricing model that cuts inference costs by at least 10 X. He co-created RWKV, the first attention free AI architecture under the Linux foundation. Eugene, you ready to take us to the top? Eugene Cheah (00:26) Yeah, when you look at this fundamental piece of technology called AI, it's something that we believe that shouldn't be controlled by only a handful of companies where they can choose to restrict your access and what you do with AI. So because of that, we fundamentally believe that people should be able to make their own choices and decisions when accessing AI models. And the best way to do it is to support all of them in the open source ecosystem. And that's why we built Federalist AI to support any AI model that you have on hugging face and to provide instant access to all of that. Nathan Latka (01:04) And so here's a list of a lot of those models for the non-technical listener that's using right now. Can you dumb this down for a second, explain it like you're explaining it to a kindergartener. Eugene Cheah (01:15) So AI models can be used for various use cases. You have the typical chat GPT use case. You can have the AI use cases for supporting people in particular language or domains. So ⁓ for example, are AI models specifically tuned for the use of providing agriculture advice for farmers in both the the Asia region and also a few models specifically for the North American regions. There are also AI models specifically Nathan Latka (01:46) Is this an example, Eugene? I'm gonna give a visual here as you describe it. So if I type agriculture here and hugging face, these some of the examples you're talking about? Eugene Cheah (01:56) Not exactly, but yeah, so for example, if you search the one, sorry, I will need to find the examples all over the head, but you can actually search more specifically like languages. So if you look on the left hand side on hugging face for languages, you can actually filter AI models by language itself as well. like something is easy for us to forget that. that the English-speaking nation is less than half of the world population. But there is actually a lot more people out there. And that's what kind of like got me to work on supporting AI models in languages and use cases beyond the mainstream. Like for example, my grandma, she speaks seven languages. She lives in Malaysia, but English nor Chinese is neither one of them. And if you follow some of the mainstream models, she wouldn't be able to speak to AI. And that's what I fundamentally want to prevent from happening, to ensure everyone in the world have access to AI. Nathan Latka (03:02) What language does your mother speak? Eugene Cheah (03:05) So my grandma, she speaks Thai, ⁓ Cantonese, Hokkien, and a little bit of Bahasa Malay, ⁓ respectively, because she's in Malaysia, and those are her main languages. Nathan Latka (03:21) So let's use tie. Let's use tie for example here. So I'm in languages. I just went to tie and I see like there's one, there's, one called text classification, 67,900 downloads. Is this an example of sort of what you mean? Eugene Cheah (03:34) Yes, so but this is a text classification model. You may want to actually use more of a text model itself. ⁓ So for example, one of popular models in Asia is the sea lion model, ⁓ or SEA. SEA, yeah. So ⁓ you may need to remove the language filters, the reset filters, yeah. So this is for example one of the lines of model that supports most of the major Southeast Asian languages. Nathan Latka (04:14) So give me an example of how your mom would use this or how this enables someone to build a model or a version of ChatGPT maybe that your mom could use. Eugene Cheah (04:24) So they could use existing applications or even go to Feather's AI inbuilt application as well. We have a chat application at the top there. And they could actually assess one of the many models ⁓ that we have in our catalog. And this includes some of models that we have highlighted respectively. Usually what most people will, what we see from the community is that like be it through the redates or local communities that people will actually find their own preferred model. And they'll know the models that they would like to run and they will actually come to our platform and they'll make the request and we'll add support for it and then they can run it straight away. Nathan Latka (05:05) So let's use the, it's sorted right now downloads high to low. Is this pronounced quen? Quen ranks number one in terms of number of downloads. Eugene Cheah (05:13) Yeah, this is one of the most popular downloaded models. So when the user signs up on the platform, they can have access to any of these models. So Quen in particular is ⁓ one of the latest ⁓ models made by Alibaba. And it's ⁓ strong in both English and Chinese and a few other ⁓ traditional languages. Nathan Latka (05:34) You guys can read the description of sort of what it does up here. So let's keep going down the stack here, Eugene. My goal on this interview is to help take a very technical sort of concept and help my listeners understand why will you build up featherless.ai is so important in terms of the end value, right? So are mostly, are, it mostly developers that are paying you for featherless? Eugene Cheah (05:53) Yeah, so developers, we also see an increased wave of what I call prosumers, and to us who's like not exactly purely developers. Some of them would be like coming in and it's like, I want to run my cloud code or I'm white coding my apps and I want my apps run on AI. So these are the two major categories. Traditional developers are building apps that uses AI and these groups are and to us as well. Yeah. Nathan Latka (06:19) And you have multiple different price points here, but what would you say the average customer is paying you per month today? Eugene Cheah (06:25) So the average entry level customer is paying the $25 a month plan. That provides them access to any of the model where they can make unlimited requests, limited to one request at a time. And then subsequently, once they figure out which models they want, and sometimes when they create apps and ship it to the app store, that's where they go to our scale up plan. And that's where they have much larger dedicated capacity because you don't really want your production apps to Nathan Latka (06:50) Who's paying for the credits though there? If they pay you 25 a month and then they use some of the models that you help them sort of play with and then it gets used a bunch, who's paying for their credit? Surely they're paying for their credits. They're putting in their API key, I guess, and they're paying for the credits? Eugene Cheah (07:03) Yeah, so we provide the entire infrastructure and the infant service. So the $25 a month is paid directly to us. And then subsequently, they are able to access the growing catalog of 30,000 models, which we are scaling respectively. Nathan Latka (07:18) How do you manage the margins if they're using a model that charges you like a dollar per API call but they're only paying you $25 a month? Don't you go underwater on their plan? You're losing money on their plan once they do 25 API calls? Eugene Cheah (07:30) So to be clear, we don't serve OpenAI or Anthropics closed source models. We are hosting the catalog of open models. Nathan Latka (07:40) only open source. So they don't charge per credit, these open source models. Eugene Cheah (07:41) Correct. ⁓ Correct. And these are models that you could have downloaded on your own high-end GPU if you have the GPU and run it. But in our case, we will just load it onto our servers when you make that request. So essentially, that $25 a month is you're getting a fraction of the GPU compute ⁓ that you have access to, and then we'll run the models on it. Nathan Latka (07:47) Got it. This makes sense. This is like back in the days when everyone went out and bought server racks and put it in their closet, right? Versus now just using AWS. You're sort of the modern day version of that for AI open source models. Eugene Cheah (08:22) Correct. So for example, like how Heroku or Versailles abstract the whole infrastructure layer for just deploying applications, we are abstracting that for AI models. And if you want to host and run at larger scale, we also provide the pricing plans for it. Nathan Latka (08:38) I see, very cool. Okay, it's now making sense to me and I'm not a tech person, so I imagine my audience is now following along nicely. How many customers are you serving now today? And then we'll get your backstory here. Eugene Cheah (08:49) Yeah, so we are serving around 10,000 plus customers and I would jokingly say the average customers do not know what a B200, MI325, they may have heard of H100 because it was in the news. And that's essentially what our job is. We abstract away all the complexity of running these AI models ⁓ and the infrastructure for it. ⁓ and these growing collection of open models include some of the best models that are already on par or surpassed with let's say Cloud Sonnet or even JPT for our mini. And we realized actually for a lot of customers when they move to production, these are the models that are more than sufficient for their needs. Nathan Latka (09:35) Interesting. And just to be clear, if you've got 10,000 customers paying that average monthly price point, you just told me of 25 a month. Eugene, can I multiply those together to get your monthly, you know, your monthly revenue is about 250,000 a month. Is that accurate? Eugene Cheah (09:49) No, it's higher than that because those will be for more towards individual users. But for the scaled up customers and for the scaled up customers, which actually represents more like half of our revenue, these customers will be coming in and saying, hey, I would like to have 100 concurrent requests. I would like to have 1,000 concurrent requests, which is a much higher volume than these individual users. And this is a pattern that we see increasingly where where developers will come in, test, prototype, find something in like, and scale up. Nathan Latka (10:23) Eugene, I'm super impressed. So sorry, just to be clear, I don't want to put you on the spot. I didn't prepare you for this, but are you comfortable sharing your monthly revenue today? Is it more like $500,000 a month? Eugene Cheah (10:33) ⁓ That's something that we are scaling up towards ⁓ and we are not yet at that number. these are those things that we already like right now the inference market has been growing rapidly that some enterprise customers that we are talking to, we are negotiating like multi-million dollar contracts on an annual basis. So it's something that we can probably hit within the next few months. Nathan Latka (10:58) That's exciting. just to be clear, we're recording in April of 2026. You're doing more than $250,000 a month, but less than $500,000 a month, but you're growing so quickly. You think you'll pass $500,000 a month sometime in fall here of 2026. Eugene Cheah (11:14) Yeah, AI is really exciting, especially in this space where some companies literally are renting entire clusters of AI servers to automate their workflows. just renting even two of these biggest servers at full capacity, let's say the B300 or the MI325 for the AMD side, are already with a million dollar contract per year. Nathan Latka (11:42) Are you holding this inventory on your balance sheet? I mean, are you running a physical plant holding all these things somewhere or are you renting these from somebody else? Eugene Cheah (11:54) So for the vast majority of our inventory, we do rent it from someone else. ⁓ So we rent it directly from the data center. But once again, I say, our average customers, do not know how to set this up or configure or manage it. And for the data center providers, they don't really want to take the responsibility of standing up the entire software infrastructure stack. And that's where, for some of our private cloud or dedicated customers, We handle all of that for them. The public cloud that you see is just one way for the customers to try and experiment. But when they scale it up, we can either provide from our inventory or even use the customer's inventory. Nathan Latka (12:37) Eugene, tell me more of your backstory here. We need to write the first line of code for Featherless. What year? Eugene Cheah (12:43) So, Fedulus started out kind of like as an accident that outgrew itself two years ago. So, as you heard, ⁓ we started from the RWKB committee where we were doing experiments in next generation foundation models. And that's where our roots are. Like this new AI architecture has the potential of reducing inference cost by over a thousand X. And if you hear all the energy demands, that is extremely lucrative. The downside is that this new generation of AI architecture hasn't been proven and scaled to the largest model size. so, correct. So multiple members of the team was working on this project. And so during the early stages of the inference platform that we did, well, this shows a lot of potential and we are still scaling it to eventually replace all the AI models you see today with something much cheaper. We created a platform where Nathan Latka (13:20) That's this, just to be clear, right? Eugene Cheah (13:41) you could fine tune our models for specific tasks. And this is what people use small models for. Because if you use small model, you have to fine tune for a specific use case. We end up having so many people fine tuning on our platform that we needed to create a platform, inference platform that can support thousands of models. And then when we end up creating that, we're like, why are we only running RWKB here? So we added support for all the other models. And then suddenly, All the signups happened and we are like, we used to say on our website, featherless, Instant access to thousands of models ⁓ made cheaper and powered by RWKV because there's some truth to it. Like underneath the hood, we use RWKV to optimize the inference, but we actually realized it lowered the conversion rate. We realized people just want, I want this model, I saw your price tag, that's all I need to know. Nathan Latka (14:32) That's amazing. So you can see our RWKV here inside of now Featherless. But what you're saying is you effectively were doing all the research here. You built this Featherless for yourself and you said, my gosh, everybody else needs this as well. What was the first massive sort of signup surge that you saw? Was it a article on, on hacker news or somewhere else? What got, or Reddit, what got you the first bump of signups? Eugene Cheah (14:56) So when we did the first soft launch, we actually just, we had a few members of the committee just posted on Reddit essentially, to some of the existing AI committees, ⁓ because we knew that we wanted to serve the models that no one else hosted. And that's what really drove the traffic. ⁓ You see most providers, they only provide, let's say, less than 100 models. That covers 50 % of our inference workload. It's the bottom 50 % where they run all these interesting fine-tuned models that people came on board for. And they are all usually very unique use cases or languages. Nathan Latka (15:37) Is this you? Is this your silly tavern AI? that you? Eugene Cheah (15:41) That is one of our members that probably did the initial post. We have three founders. ⁓ So this is probably it was done by Wes. So Wes, Wesley George was our CEO. He's based in Toronto and Harrison Vandevelle. He's based in Australia, our CTO. Nathan Latka (15:47) co-founders. How many co-founders do you have? You Okay, wow, so 2024 was official launched in two years ago. And you're doing more than $250,000 a month today in revenue. So it's fair to say you've gone from zero to a million dollars of revenue. What like very quickly, right? In a couple of months. Eugene Cheah (16:19) Mm, mm, yes. So, so. Nathan Latka (16:22) You're like, you're an engineer and I'm a business guy. So it's uncomfortable when I ask you finance questions, but I love to capture the growth story. Eugene Cheah (16:30) You want to hear the funniest bit about this right? The company at that point of time was called Recursor. So we already had an inference platform for RWBKB. Recursive model, RWBKB. It all makes sense. And Federalist was meant to be a pricing experiment. So we gave it a different name. But within the first few days, it became more profitable and more revenue than the original company platform that we were like... Nathan Latka (16:34) Yes. Okay. Eugene Cheah (16:59) I guess we are featherless now. Nathan Latka (17:03) That's amazing. what did it say? mean, did you guys go, I mean, you guys, as your co-founders, you must have said, oh my gosh, we just passed 88, $83,000 a month in revenue and we've only been live for like three months. How many months did it take you to break 83,000 a month? Do you remember? Eugene Cheah (17:18) wow, I can't really remember that moment but it was like... It was all such ⁓ a blaze because like it was a case of like we get more users, the servers are on fire, we add more servers, we get more users, we add more servers and ⁓ so it was like just a constant hectic rush there. And it wasn't until like ⁓ much more recently where we had a lot... we recently... ⁓ had a close our funding our latest a round where we had where we had it now like enough server capacity is like we have a sign already now the there is slightly more servers and users for now Nathan Latka (18:01) for now. That's great. Tell me tell me more about the funding history. You just said you just closed a series a when did you close that and how much was it for Eugene Cheah (18:10) So we closed our series A ⁓ sometime during December. That was a $20 million round. ⁓ And the leads for it was Airbus Ventures and AMD Ventures. And since then, we've been scaling the platform much more aggressively. Nathan Latka (18:28) So you closed 30 million series a in December of 2025, 20, 20 million. Okay. And what about the seed? Did you do a seed round or did you skip it? Eugene Cheah (18:31) No, 20. Yeah. We did do a seed round much earlier and ⁓ that was pre-Fedellus. That was when we did a seed round over the research of RWKB. Nathan Latka (18:48) and how much was the seed round for? Eugene Cheah (18:51) The seed round was a small 2 million round. Nathan Latka (18:55) Okay, this is great. How much would I have to bribe you to let me put in 100K on the Series A round and I'll promote the hell out of this to drive you guys more users. Do I have to bribe you? Do you take bribes? Eugene Cheah (19:06) Ha We accept users. Nathan Latka (19:11) ⁓ So there's zero chance that I can write a 100k check on the Series A docs. I won't even argue about valuation. I won't even do any diligence. I close my eyes, I write a 100k check, and you and your co-founders take it. Eugene Cheah (19:26) ⁓ I mean we do have some flexibility to do a safe note here and there as part of this round but yeah, as long ⁓ as you can help us get users, think our investors are more than happy to entertain the idea. Nathan Latka (19:41) Well, I'm a marketing guy and we've invested 250 million and 700 portfolio companies that I think should all be looking at what you're building. So I'll take this offline with you to try and convince you to let me put in 100k on a safe note. You can make up whatever valuation you want because I think this is, I don't understand everything you're doing. I'm not a technologist, but I love how you talk about building the product and the growth speaks for itself. Eugene Cheah (20:03) I'm all ears for that and I would like to know how many of your portfolio is burning so much AI usage that we can come in to help them lower their costs. Nathan Latka (20:12) Tons. mean, this is why it's interesting, right? When we look at all of the profit and loss. So every portfolio company has to connect their profit and loss to founder path. I can look in their cost of goods sold line and I see how much they're paying to Anthropic and open AI and on all the credit spend. If you're telling me that you've got a way to help them cut their cogs in half, right? Or even more, that's extremely valuable. Eugene Cheah (20:37) Exactly, and that's actually how our sales team are starting to close a lot of these deals because they'll come in and say, hey, what are you using AI for? Do you know for half of this workload, you could use this open model that's much cheaper and lower? We can provide that. For this half of this work, you can use this model. And since we have seen them all, we can advise more specifically. Nathan Latka (20:59) I love that. how much of what you do, would you say is sort of note people buying you sort of without emailing you without a call versus high touch you telling people what model they should use or how they can save 10 times or, 10 X their inference spend. Eugene Cheah (21:12) It's it's currently a bit of a both. So when it comes to the individual users, when they're coming in on the public cloud, this is true word of mouth self discovery for most of the cases. And every now and then some of these users will upgrade down the path respectively. But we also realize that there is a lot of money on the table right now where you can go after the startups that, hey, I just built my entire startup or SMB. on OpenAI or Entropic and I'm $100,000 a month and ⁓ I do not know what I was doing. And it's like, okay, we can help you here. We can lower your bill by half and we can see where it goes. That is usually some of most ideal large volume customers ⁓ when they are basically spending that much and they are entering a situation where, hey, we need to start thinking about this and then we step in. Nathan Latka (21:50) Yep. Yes. Okay, so here's a deal. I know you're an engineer, you're not a deal guy, but I'm gonna try and sell you on a deal. If I can bring you five customers that are spending more than 100 grand a month paying for credits for the models and they sign up with Federalist, you'll let me put $100,000 in on a safe note. I don't care what the valuation is. Do we have a deal? Eugene Cheah (22:26) I think I sometimes have a hard time. Nathan Latka (22:31) I love this show. is great. Tell me more about your team, Eugene. How many people are full time today? Eugene Cheah (22:40) So we currently have around 27. We are close to past the 30 mark soon. We are aggressively hiring. ⁓ And our team is split across both the platform for deploy engineers, which supports the customers, the infrastructure team that keeps everything running but doesn't do anything with building, for example. And then we also have the research team that is still working on the other RWKV line of models because we still feel that there's a lot of room still to optimize these AI models today. Nathan Latka (23:14) how many of the 27 are working on infrastructure. Eugene Cheah (23:19) ⁓ So ⁓ around 12 is working on infrastructure right now. ⁓ And then 10 is moving towards platform, go to market. And then the ⁓ rest is ⁓ increasingly like GTM, like marketing activities and research. Nathan Latka (23:25) Okay. Okay, got it. So 12 on infra 10 on like platform and then the rest, you know, call it five, six people are sort of research and go to market motion. Okay. Eugene Cheah (23:50) Yes, the platform does do some marketing as well. They do more like like dev rel and things like that. Nathan Latka (23:58) And is all of the inbound right now pure word of mouth or just Reddit or how are the customers finding you? Eugene Cheah (24:06) ⁓ Both via Reddit, Huggingface, various other platforms. We have started trying to, we also have started like preparing marketing materials, but those haven't really kicked off yet. So it's mostly been Reddit and what not. Nathan Latka (24:22) Where do people find, like I'm on hugging face right now. Where could I click on hugging face to see, hugging face promoting featherless. Eugene Cheah (24:28) So if you go to Models, so this is a little bit tricky because it's, and you go under Others, oh there, you can see it appear. There's a pop-up there. You can see Federalist AI. You can filter the inference provider. Federalist AI. Yeah, so you select Federalist AI on the left-hand side. Nathan Latka (24:36) under Where? Filter by inference available. Featherless, why don't I see featherless? Eugene Cheah (24:55) You just went over it. Yeah. Yeah. Yeah. So you reset all your inference provider filter and select only $5. So what happens is when you come to some of these models, so for example, you can look at the... Let's take the recent one. That starts with an nand bridge because Quen is quite popular. A lot of providers are providing it. But if you go down the list slightly, Nathan Latka (24:56) here, here. Eugene Cheah (25:23) One of the step fun models, for example. Nathan Latka (25:28) You're saying this this step the what the step what? Eugene Cheah (25:31) Step fun, yeah that model. For example, if you click on that, ⁓ we are one of the only providers online right now for this model. Nathan Latka (25:33) right here. ⁓ so this is the pop-up here that you're getting promoted. Eugene Cheah (25:47) Yeah. Yeah. So if you come, if you, if, if you are paying card game fees on the pro accounts, you can actually instantly do inference on fair dollars for these models. So I'll, I'll go to market strategy competition here, right? If you're not going to fight the giant battle of like the top 10 models. Quinn is one of the the top 10 and top 100 models where all the providers are fighting it out there. You're going to see like eight, 10 providers. Our strategy is Nathan Latka (25:49) I see. Eugene Cheah (26:16) We are supporting all the other models that people are interested in experimenting. And for example, the step-1 model is a particularly popular model for us that easily shipped several contracts for us on this model alone. And no one else is doing it. Nathan Latka (26:33) Why are you guys, I mean, I know you're smart, right? I I don't understand all the research, but I can see you're doing a ton of research, but why are you the only infants provider listed here? Is it just really hard to build that inference model? Eugene Cheah (26:47) Yeah, so right now today traditionally when you do inference as an infrastructure provider, when you load a model onto GPU and set it up, it takes around 30 minutes, give and take, for the models to load and get up and running. That 30 minutes is cost time. When we're talking about GPUs that cost, let's say, even $5 an hour or an entire server that can cost $80 an hour. You don't want to half an hour here, half an hour there, letting it run just to load a model. So operationally, as an infrastructure company, you do not, an inference company, you do not want to be switching out models left and right to cater to various user demands, because you're going to waste all this switching time. We built our inference stack and our infrastructure design from scratch where we load models in. 5 to 10 seconds or 30 seconds and under for the especially large one. And what this lets us do means we can actually change our GPU to serve the various demands for various models live. And that allows us to host a lot more models. When we say we support 100,000 models, right now 30,000, it doesn't mean we have 100,000 GPUs all the models live right now. It just means we have, let's say, a few thousand GPUs. hopefully a lot more as we have a lot more customers. And when they ask for a model, we load it into the GPU on demand. Nathan Latka (28:24) And that on demand is what makes you unique and the only reason you can do that is because you built your whole stack from scratch. Eugene Cheah (28:32) Correct. We heavily customized our stack from the get-go to specialize into being able to swap and load these models dynamically. Nathan Latka (28:41) So is it fair to say then if I go back, mean are these your biggest competitors over here? Eugene Cheah (28:50) Yeah, when you talk about the top 10 or top 100 models, yes they are. If you talk about the rest, that's where we come in. And in an AI landscape where, try to view it the other way. Like today, a lot of companies have started fine tuning their own models for their own unique use cases and specialization. ⁓ In a world where various companies are fine tuning their own model. You can't go to a provider that can only support 100 models. There are more than 100 companies on earth. You want an infrastructure tailored to be able to handle all these various fine-tunes. Nathan Latka (29:30) So do you have the largest coverage? mean, is that what you measure? How many models can you cover? Eugene Cheah (29:36) Yes and that also allows us to have a lot of demand for all these models like like step fun for example is not an unpopular model it's shipping billions of tokens per day. Nathan Latka (29:50) How do you know that? Can I see that somewhere on Hugging Face? Eugene Cheah (29:55) Yeah, unfortunately, I don't think Hanging Face provides that statistics, but if you go by the download count, it's quite a popular model. You have to understand that this is a 200 billion parameter model, meaning you need at least some of highest end GPU. We're talking about at least four H100. We're talking about $20 per hour systems to run this model. And we're talking about hundreds of thousands of companies are already using this model itself. Nathan Latka (30:25) And give me an example just so my audience understands. mean, can you, like, do you know the physical location of the data center that you're renting space from? Eugene Cheah (30:33) Yes, ⁓ so we primarily ran from ⁓ North American data centers where we from a few key partners, we are increasingly also purchasing capacity in Europe from a few key partners and providers ⁓ respectively ⁓ because ⁓ for the sovereign AI demand in Europe, we are actually getting more and more customers that say, I only want my AI to be within a certain ⁓ sovereignty and we are actually helping facilitate that as well. Those are our key major regions. Nathan Latka (31:05) And so it sounds like your biggest monthly expense, maybe you're spending north of a hundred thousand dollars a month on renting these different data center spaces. Is that, that accurate? Is that one of the reasons you needed to raise capital so you could go buy space faster? Eugene Cheah (31:21) Yes, and also one of the things is to actually to get direct access to and from the suppliers as well because I think that's one of the most exciting thing about, would call it a happy accident that where we got investment from AMD was that if you kept track of the GPU space or even compute space in general, in the past six months, we went straight to inventory being wiped out. Not just on the GPU side, on the CPU side, you may have heard within the consumers, people are complaining about RAM prices shooting up through the roof. Supply has been heavily eaten up, but by raising capital from strategic investors like AMD in particular, and a few other who manufacture some of the GPUs and the systems as well, and ⁓ Airbus Ventures and a few other Nathan Latka (32:10) Like this one here, right? AMD. Eugene Cheah (32:21) ⁓ investors in the industry, we are able to get access to the supply without going through all these hurdles right now. So that has been growing advantage for us as well. Nathan Latka (32:34) Interesting, got it. Yeah, so you've picked investors that can also help you manage your cost and accessibility to this sort of data center processing that you need. Yeah, interesting. Very, very cool. Okay, I've learned a lot on this episode. I guess let me just ask one or two other questions. I know Coheer just from my background in the SaaS space. I don't think there are technologists like you. Why do they have an inference provider option over here? Eugene Cheah (32:45) Correct. Exactly. So Cohere in particular, ⁓ they have their own particular line of models that they created ⁓ for basically the North American market and in particular Canada. So they are going to the direction of highly tailored sovereign AI models ⁓ for the domestic market. And we actually see this happening more and more. So for Cohere, they will service the Canadian market. The US market is going to be served by...

Data and Sources

All figures on this page are taken directly from interviews or are estimates from public sources and proprietary models. Not financial advice. Read full disclaimer.

Claim this profile