Embeddings index, what is it exactly?

Operating system

macOS

Joplin version

3.7.18

What issue do you have?

Hello,
Discussions about AI are important and I believe it's a good thing to have such features in Joplin, disabled by default, with good privacy and security.

I just wanted to clarify what are exactly embeddings and index of embeddings in the bottom of the AI settings. I have seen it's disabled by default and i didn't enable it even if i already use the AI feature. I think it would be nice to explain what it does exactly when enabled.

I imagine it creates an index of embeddings of all our notes but is it sent to cloud providers when enabled ? only the current note ? all the notes ? is it used only to help local search ? what are embeddings exactly ? is it safe to send embeddings of our note to cloud providers ?

Can someone please explain exactly the usage of this feature ?
Thanks

I have found that it's actually explained pretty well in the docs, albeit under the heading of "Semantic Search" (which is what the embeddings let you do)

Ok thanks for the link
And a small question about that, if i use this local embedding index, is the content of this index sent to Cloud AI providers if they are enabled ?

no, the embedding is all done locally and the resulting index is stored locally (it doesn't even sync - each device builds its own index).

All of this runs entirely on your device. The model is local; no note content is sent to a cloud service. The index is also local β€” it is not synced β€” so each device builds its own.

Just to expand upon "What are embedding exactly" since the rest seems pretty well answered.

If I mention the word orange out of context, you might associate it with various things. Orange is a colour. Orange is a fruit. Strangely, the fruit also seems to kind of relate to the colour. But also, Orange is a telephone mobile network (or was, anyway).

We could begin to say then, that there's a higher association between the word orange being a colour, which might be 95% of how I'd probably use the word orange, maybe 20% of the time I mean the fruits, and maybe 0.001% of the time I mean the mobile network.

So if broadly I ask the computer, "What invoices have I stored recently for my bills", and I have three notes by way of example, one a review about the film A Clockwork Orange, the other a recipe book, and the third a collection of monthly bills, the computer might be able to determine that the Orange network is a telecom bill and present that note, while avoiding bringing up recipes for fruit salads, because the embeddings are what allows for that kind of associativity on a probability level. When it gets more advanced, words can pair up with other words, so the computer might know "Clockwork Orange" refers to a film more than it's likely to refer to a literal orange constructed out of iron and motors, and that context begins to shift probability in more basic ways than me being vague above saying I mean the colour 95% of the time.

By itself the embeddings don't actually do anything and are just an abstract series of numbers that humans wouldn't find very intuitive, but they're what's needed for computers to begin understanding more complex associations in the case above for semantic search (and others) to function efficiently, by precomputing a lot of this association so that it doesn't have to be done over and over for efficiency. The semantic search then can operate on this embeddings index to be speedy, but you could perhaps have other smart stuff go on, like say automatic tagging of related notes by subjects, people in them, & etc.

So in simple terms, an embeddings index is a specific type of database that allows for the computer to understand deeper concepts behind language rather than just raw text; more mathematically, it's a collection of vectors that associates words to other words in super duper high dimensional space.

Great explanation, thanks.
I have also discovered while searching a bit about that that embeddings could be reversed to try to get real text out of them even if they are just vectors. That means that they still are quite precious information and it's a good thing they remain local :slight_smile:

Yes,

If you put your password into Joplin, and allowed embeddings, then the embedding index itself would potentially store a separate copy of your password and might start making vague associations that you always seem to enter "hunter2" around "Gmail". (Not necessarily as easy as I make it sound).

Course, you shouldn't store passwords in Joplin anyway, but it demonstrates some of the potential concerns. While you can't use the embeddings to return the original notes, it could still potentially leak context about what's actually in them, hence why it doesn't get synced and is local only.

More generally, If I was constantly writing about particular perhaps taboo subjects, like Nintendo emulation, I use Arch btw, and workplace unions, the embeddings technically could suggest that I'm inclined a certain way if someone managed to grab the index, even without having the raw notes. It's super unlikely realistically, and again, the embeddings never leave the device.

Thanks again.

Right now my embeddings are created but I can't see if a search result has been made with the current words or with the semantic search. Is there a new to know that ? Semantic search are enabled in my settings but how to be sure which type of search managed to get the results ?

I think currently results are either basic search or semantic search but not both, per

I'm not sure if the UI might change to make it more clear in the future which was used, or to be able force the semantic search even if regular searching has matches or etc, but it's possible that feature will evolve with feedback as everything else does.

I think currently results are either basic search or semantic search but not both, per

Related notes:


  1. This might change in the future. See the related configuration, where the k is the maximum number of results). β†©οΈŽ

  2. Related code: SearchEngine.ts and SearchEngineUtils.ts β†©οΈŽ

  3. Related code: canSemanticSearch_ in SearchEngine.ts β†©οΈŽ

Thanks for these details.
Is there a way to know whether semantic search is enabled or not ?

I have enabled it and here is what i see in the settings :

It says Indexing is waiting and the number of notes is larger than the total of notes, is there an issue ? does the semantic search work when the indexer is in this status ?

Thanks

It should work. To test I suggest making a search with keywords that you know are not present in your note collection - it should normally return 0 result, but with semantic search it might return some.

Another good test is to make a search in English for notes that you know are in French, and it should return those French notes (which is a good indication that it search by meaning rather than by exact keywords)

Thanks for your answer.

About the indexing, I have seen that "En attente" in french is "Idle" in english which means it should be available. Only small issue about the number of notes indexed which is greater than total number, that's strange. Maybe because I have deleted some notes ?

An maybe it would be interesting to have a way to delete the index in the AI settings to be able to delete and stop using it, or delete to make Joplin recalculate it.

I have tried to look at english words when i have french notes and it doesn't seem to work. I have no example where it works.

Only example I managed to make work with semantic (i believe) is with Joplinn on a very small note

Maybe my notes have too much text to work correctly .. do you have an easy test i could do ?
I will create small notes with only some sentences to see if it may work with them.