p.enthalabs

Show HN: SubSmith – Turn your own videos into language-learning material

subsmith.app · Read Story HN original

I've been learning Japanese for a few years and kept running into a similar problem. I'd find a video I wanted to learn from, hear a useful sentence, and then realise that turning that sentence into something I could study later was both time consuming and draining at times.

I would end up jumping between a video player, subtitles/transcription, a dictionary, screenshots, audio clips and Anki. So I built SubSmith to bring that workflow together.

You can drop a video or audio file into it, generate a transcript locally and then use the transcript alongside the media to:

* look up words and sentences * replay individual lines * edit the transcript * save useful sentences with their original context/audio * export them as Anki cards

The important part for me is that it works with your own media. It isn't tied to a particular streaming service or library, so I can use the random anime episode, podcast, lecture, etc. that I'm actually interested in studying.

It's an offline-first desktop app, and transcription happens locally rather than sending the media to a transcription API.

I'm sharing it here because I'm now more interested in finding out where this workflow breaks down for other people rather than adding features randomly now that I have solid core/base.

For example:

* Would you actually save sentences from your own media? * Which part of this process feels like too much work? * Does having the audio/context attached make creating an Anki card more useful? * Would you prefer this to work inside your existing video player/browser? * Is installing a desktop app a significant barrier? * And does requiring an account before starting the free trial make you give up?

The current version does require an account to start the trial, and I'm trying to work out whether that's meaningful friction for the people who would actually use this.

It's free to try, and I'd particularly appreciate feedback from people who already learn languages through their own videos, anime, films, podcasts or other media.

I'm the developer, so I'll be around in the comments to answer questions and discuss how it works.

https://subsmith.app

Comments

Cool idea!

While I'm not so sure about learning a language with content from fictional sources, that should at least make it more relevant for the learners.

And you can use it on podcasts too, where people talk normally.

The biggest hurdle is probably that users need to get hold of the media so they can use it as input.

Thanks! There is a whole like subsection of language learning called immersion based learning. Fictional sources can be OK just depends on how specialised and unique the vocabulary is.

Yeah that is my suspicion too, I am considering adding an option to transcribe system audio too so it could hook into whatever video or audio is playing on the computer. That would be quite a big change so would like feedback first! hahah

I love having auto-generated subtitles available, but I am still skeptical that we can call them "accurate" yet. For many languages the models are quite weak, and even for English and German on YouTube I see errors all the time.
Yeah I agree the Youtube ones are not very good, I am not sure what they have implemented behind the scenes for it. However this uses the Whisper Large model which is quite a bit better than the Youtube auto generated ones. Only catch is you need to have the video on disk and wait a bit longer lol
Youtube is very bad, but I have honestly not yet seen anyone implement this well on any platform, which indicates to me that it's not a solved problem. Very open to being proven wrong here though - do you have examples of Whisper Large in use in a large consumable sample?
Even for English (which surely has the best support) anything outside of urban American accents is still incredibly weak - you can pretty much guarantee very noticeable mistakes on even highly intelligible mainstream urban British/Irish/Aussie/Kiwi/SA accents. Stronger regional / rural accents & dialects usually render generated subtitles entirely useless.

Definitely a domain a million miles from being solved if even English is this bad.

I've been learning Japanese for just over a year now, and its a community that might have the highest concentration of amazing free learning apps with sentence mining capabilities, many of which do things very similarly to what you're building. Have you tried some of these?

asbplayer Manatan Yomine Anki Miner AnimeCards Nagare

These are just for video, there are other similar solutions for games, manga and ebooks, and some of them are all in one

I've only heard of asbplayer and Anki miner before, the others I haven't. It's a great space tbh, people are building all the time. For me specifically it's about trying to serve all languages to the best of my ability, alot of the ones I've come across only focus on Japanese or a small selection of languages. Also I think there's scope to improve UX and reducing friction for non tech savvy people
even for "tech savvy" people, it's great to have an easier to use UX, or at least have the option to
In my experience, apps for learning japanese is a hugely saturated niche because the venn diagram of people learning japanese and people who like building software is a circle
I build the exact same system, but for (mandarin, simplified) Chinese which i'm learning. Wanted to create such a system for a while, but it's a lot of hooking up the correct libs and making a UI etc. never found the energy. Thanks to 'ai' I was able to do this finally, think more people needed this push :-) It's web-based and can be self-hosted in a docker.
Care to share with a fellow learner? :)
Nice! Yeah I think the functionality is never very complicated, I'm more so interested the UX and trying to add features that reduce friction and slowly build an app that can be stable and relied upon.
In the age of LLMs, AI assisted programing and vibe coding can't this all be replicated?

I'm confused why is this a yearly subscription or paid at all?

This tool might as well be open source to be brutally honest.

Yeah that is a fair point, it can be replicated but not many people want to replicate an app and would rather pay for it. The functionality is not hard to replicate but the UX side of things and general support and improvements over time is what people have paid for. There is a lifetime option also - yearly was about offering choices.
Same goes for Dropbox - you can already build such a system yourself quite trivially by getting an FTP account, mounting it locally with curlftpfs, and then using SVN or CSV on the mounted filesystem.
I understand your jest, but you can really vibecode your own Dropbox in this day and age.
I've been building something similar but targeted at bringing content that interests you to your level: https://katarineko.com.

I've been learning Japanese as well and the content at my level is boring while the content I'm interested in is a wall.

Ahh nice, hope it goes well!
How does this product compare to the methodology of dreaming.com?
>Drop any file. Get accurate subtitles.

This is my main issue with tools like these. Automated subtitles are inherently inaccurate, because much of language depends on context which you cannot get just from a brief snippet of something. This wouldn't be much of an issue if the tool was targeted at people dedicated to creating subtitles, but if you're trying to learn a new language, how are you supposed to know when the output is wrong or not?

Yeah that is a very good point, it’s something I thought about a lot when positioning the app. Essentially there are 2 mitigations to this, one is that it’s not really ideal for beginners and mainly targeting intermediate learners who can recognise or at least question errors to a certain extent but don’t want to create the entire subtitle file themselves. There is an inline editor for them to fix the errors.

The other thing is the speech recognition models are very sensitive to the audio. If the audio quality is good then the model would do a good job, I’m not sure how much experience you have with Whisper Large but it is very capable on normal speech. The issues arise when there are many competing sounds overriding each other

That’s fair enough! I am curious tho, are there any issues you face while using the app that could be handled better? Or any features that would make it easier to use?

ASR in LR is down often from my experience.

I can't even login to languagereactor.
Super cool, great job!
account required != free to try.
I've been somewhat studying Japanese for... quite a few years now. I have yet to find the "ideal" method, but I know it must be mobile (I'm continually migrating away from sitting in front of my laptop) and something I can randomly get in and out of at any time (anything taking up dedicated batches of time eventually gets dropped or barely touched).
> I'd find a video I wanted to learn from, hear a useful sentence, and then realise that turning that sentence into something I could study later was both time consuming and draining at times.

How does not study a sentence? You mean memorizing it later?

By study I mean adding to Anki for example as a flashcard
This is really cool idea. But don't you think taking material and saving specific sentences may add unnecessary mental pressure to the process of learning which is said is ideal when the student is in a playful mood?