The interview focused mainly on ML system design, scalability, production systems, and model monitoring.
Design an architecture for serving an AI/ML model as a microservice.
How would you handle authentication and authorization for an ML inference API?
How would you handle request filtering, throttling, and abuse protection for the inference service?
How would you structure the inference microservice?
How would you handle heavy or batch inference requests?
How would you monitor and maintain an ML service in production?
How would you deploy an ML microservice?
If you have limited resources, how would you decide between horizontal scaling and vertical scaling?
What factors would you consider when choosing between horizontal and vertical scaling?
After 3 months in production, if customers complain that search or recommendation results are no longer relevant, how would you diagnose and fix the issue?
Follow-up areas included:
The round focused on embeddings, vector databases, similarity search, ANN algorithms, and optimizing vector search performance.
How are embeddings generated from text, images, or other input data?
What is an embedding, and how does a model like BERT/CLIP/Sentence Transformer generate it?
How does vector similarity search work in a vector database?
How would you find the closest vector to a given query vector in a vector database?
What similarity/distance metrics can be used to compare vectors?
How does a vector database efficiently find the Top-K most similar vectors?
What is Approximate Nearest Neighbor (ANN) search, and why is it used?
How would you speed up vector search when the database contains a very large number of vectors?
What are HNSW, IVF, and PQ? How do they help optimize vector search?
How does HNSW work for finding nearest vectors?
How does IVF (Inverted File Index) work in vector search?
Suppose you have 1 million vectors. How would you use clustering to make nearest-neighbor search faster?
How would you implement a hierarchical/two-level clustering approach for vector search?
What is the difference between exact nearest-neighbor search and approximate nearest-neighbor search?
If you have 3D vectors and need to find the closest vector, how would you solve it?
How would you optimize the above approach when the number of vectors becomes very large?
The interview started around 5:50 AM in the morning. The interviewer was Japanese, so there was a slight language/communication challenge during the interview.
The interview was completely focused on Machine Learning and some of the projects I had worked on. There was an in-depth discussion around vector search and its optimization, which was quite challenging, especially for a fresher applying for an ML role. I would rate the overall interview difficulty around 9/10.
I would suggest preparing for in-depth ML discussions, especially around embeddings, vector databases, ANN, vector search, and optimization. Also, be prepared for questions related to API architecture and ML system design.