# Week 12: Project Search Engine

**URL:** <https://discourse.joplinapp.org/t/week-12-project-search-engine/10360>\
**Category:** GSoC - Search Engine Project\
**Created:** [31 July 2020 18:00 UTC](https://discourse.joplinapp.org/t/week-12-project-search-engine/10360 "2020-07-31T18:00:51Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![naviji](https://yyz2.discourse-cdn.com/flex028/user_avatar/discourse.joplinapp.org/naviji/32/2502_2.png) [@naviji](https://discourse.joplinapp.org/u/naviji)\
**Post date:** [31 July 2020 18:00 UTC](https://discourse.joplinapp.org/t/week-12-project-search-engine/10360/1 "2020-07-31T18:00:51Z")

</div>

Added a new filter **sourceurl**.

> [@Search Syntax Documentation](https://discourse.joplinapp.org/t/search-syntax-documentation/9110):
>
> Operator Description Example Searches within the title and body of the note foo searches for all notes with foo in title or body. foo bar searches for all notes with foo in title or body and bar in title or body. foo -bar searches for all notes with “foo” that don’t contain "bar" "foo bar" does a phrase search. any: By default all the search terms will be connected by and in the backend (notebook is the only exception). Use the any operator if you want the search terms connected by…

Refactored the PR to make it easier to understand and to avoid duplicate code. Also added more tests and fixed some highlighting problems.

Made some more progress in fuzzy search. Looking into how multiple languages could be supported.

---

<div class="post-metadata">

**Author:** ![laurent](https://avatars.discourse-cdn.com/v4/letter/l/ce7236/32.png) [@laurent](https://discourse.joplinapp.org/u/laurent)\
**Post date:** [1 August 2020 16:06 UTC](https://discourse.joplinapp.org/t/week-12-project-search-engine/10360/2 "2020-08-01T16:06:07Z")

</div>

> [@naviji](#):
>
> Made some more progress in fuzzy search. Looking into how multiple languages could be supported.

Do you have a spec on how this feature is going to work exactly? Making the feature dependent on a language seems a bit complicated as it means it should be possible to change the language from the UI. Also for those who have multiple language, it's not the best UX to have to specify the language of their search query every time.

I wonder if we could make it language agnostic instead? It could perhaps fix simple spelling mistakes so for example if I search for "fuzzi", it would return "fuzzy" too. Or if I type ”我是法人“ it would return "我是法国人" too, and this it can do without knowing anything about the language based on string similarity (I guess that's just the Levenstein distance). What do you think?

---

<div class="post-metadata">

**Author:** ![naviji](https://yyz2.discourse-cdn.com/flex028/user_avatar/discourse.joplinapp.org/naviji/32/2502_2.png) [@naviji](https://discourse.joplinapp.org/u/naviji)\
**Post date:** [2 August 2020 03:03 UTC](https://discourse.joplinapp.org/t/week-12-project-search-engine/10360/3 "2020-08-02T03:03:30Z")

</div>

Yes, language-agnostic would be the way to go.

As for the spec: [Week 11: Project Search Engine](https://discourse.joplinapp.org/t/week-11-project-search-engine/10244)  
Tell me what you think. I’ll add the language part soon.
