addressing new comments after merge

Signed-off-by: Matt Williams <m@technovangelist.com>
applied mikes comments
2023-10-15 14:17:23 -07:00 · 2023-10-14 08:29:24 -07:00 · 2023-10-12 15:57:50 -07:00 · 2023-10-12 15:34:57 -07:00 · 2023-10-12 12:56:43 -07:00 · 2023-10-12 12:52:43 -07:00
185 changed files with 15797 additions and 45852 deletions
--- a/.dockerignore
+++ b/.dockerignore
@@ -1,7 +1,8 @@
 build
 llama/build
 .venv
 .vscode
 ollama
 app
-web
+dist
 scripts
 llm/llama.cpp/ggml
 llm/llama.cpp/gguf
 .env
--- a/.gitignore
+++ b/.gitignore
@@ -5,3 +5,4 @@
 .swp
 dist
 ollama
 ggml-metal.metal
--- a/.gitmodules
+++ b/.gitmodules
@@ -0,0 +1,10 @@
 [submodule "llm/llama.cpp/ggml"]
    path = llm/llama.cpp/ggml
    url = https://github.com/ggerganov/llama.cpp.git
    ignore = dirty
    shallow = true
 [submodule "llm/llama.cpp/gguf"]
    path = llm/llama.cpp/gguf
    url = https://github.com/ggerganov/llama.cpp.git
    ignore = dirty
    shallow = true
--- a/28
+++ b/28
@@ -1,15 +1,23 @@
-FROM golang:1.20
+FROM nvidia/cuda:11.8.0-devel-ubuntu22.04
 WORKDIR /go/src/github.com/jmorganca/ollama
 COPY . .
 RUN CGO_ENABLED=1 go build -ldflags '-linkmode external -extldflags "-static"' .
-FROM alpine
+ARG TARGETARCH
 ARG GOFLAGS="'-ldflags=-w -s'"
 WORKDIR /go/src/github.com/jmorganca/ollama
 RUN apt-get update && apt-get install -y git build-essential cmake
 ADD https://dl.google.com/go/go1.21.1.linux-$TARGETARCH.tar.gz /tmp/go1.21.1.tar.gz
 RUN mkdir -p /usr/local && tar xz -C /usr/local </tmp/go1.21.1.tar.gz
 COPY . .
 ENV GOARCH=$TARGETARCH
 ENV GOFLAGS=$GOFLAGS
 RUN /usr/local/go/bin/go generate ./... \
    && /usr/local/go/bin/go build .
 FROM ubuntu:22.04
 RUN apt-get update && apt-get install -y ca-certificates
 COPY --from=0 /go/src/github.com/jmorganca/ollama/ollama /bin/ollama
 EXPOSE 11434
 ARG USER=ollama
 ARG GROUP=ollama
 RUN addgroup -g 1000 $GROUP && adduser -u 1000 -DG $GROUP $USER
 USER $USER:$GROUP
 ENTRYPOINT ["/bin/ollama"]
 ENV OLLAMA_HOST 0.0.0.0
 ENTRYPOINT ["/bin/ollama"]
 CMD ["serve"]
--- a/Dockerfile.build
+++ b/Dockerfile.build
@@ -0,0 +1,32 @@
 # centos7 amd64 dependencies
 FROM --platform=linux/amd64 nvidia/cuda:11.8.0-devel-centos7 AS base-amd64
 RUN yum install -y https://repo.ius.io/ius-release-el7.rpm centos-release-scl && \
    yum update -y && \
    yum install -y devtoolset-10-gcc devtoolset-10-gcc-c++ git236 wget
 RUN wget "https://github.com/Kitware/CMake/releases/download/v3.27.6/cmake-3.27.6-linux-x86_64.sh" -O cmake-installer.sh && chmod +x cmake-installer.sh && ./cmake-installer.sh --skip-license --prefix=/usr/local
 ENV PATH /opt/rh/devtoolset-10/root/usr/bin:$PATH
 # centos8 arm64 dependencies
 FROM --platform=linux/arm64 nvidia/cuda:11.4.3-devel-centos8 AS base-arm64
 RUN sed -i -e 's/mirrorlist/#mirrorlist/g' -e 's|#baseurl=http://mirror.centos.org|baseurl=http://vault.centos.org|g' /etc/yum.repos.d/CentOS-*
 RUN yum install -y git cmake
 FROM base-${TARGETARCH}
 ARG TARGETARCH
 ARG GOFLAGS="'-ldflags -w -s'"
 # install go
 ADD https://dl.google.com/go/go1.21.1.linux-$TARGETARCH.tar.gz /tmp/go1.21.1.tar.gz
 RUN mkdir -p /usr/local && tar xz -C /usr/local </tmp/go1.21.1.tar.gz
 # build the final binary
 WORKDIR /go/src/github.com/jmorganca/ollama
 COPY . .
 ENV GOOS=linux
 ENV GOARCH=$TARGETARCH
 ENV GOFLAGS=$GOFLAGS
 RUN /usr/local/go/bin/go generate ./... && \
    /usr/local/go/bin/go build .
--- a/README.md
+++ b/README.md
@@ -9,19 +9,27 @@
 [![Discord](https://dcbadge.vercel.app/api/server/ollama?style=flat&compact=true)](https://discord.gg/ollama)
-> Note: Ollama is in early preview. Please report any issues you find.
+Get up and running with large language models locally.
-Run, create, and share large language models (LLMs).
+### macOS
-## Download
+[Download](https://ollama.ai/download/Ollama-darwin.zip)
- [Download](https://ollama.ai/download) for macOS on Apple Silicon (Intel coming soon)
+### Linux & WSL2
- Download for Windows and Linux (coming soon)
+
- Build [from source](#building)
+```
 curl https://ollama.ai/install.sh | sh
 ```
 [Manual install instructions](https://github.com/jmorganca/ollama/blob/main/docs/linux.md)
 ### Windows
 coming soon
 ## Quickstart
-To run and chat with [Llama 2](https://ai.meta.com/llama), the new model by Meta:
+To run and chat with [Llama 2](https://ollama.ai/library/llama2):
 ```
 ollama run llama2
@@ -29,32 +37,50 @@ ollama run llama2
 ## Model library
-`ollama` includes a library of open-source models:
+Ollama supports a list of open-source models available on [ollama.ai/library](https://ollama.ai/library 'ollama model library')
-| Model                    | Parameters | Size  | Download                    |
+Here are some example open-source models that can be downloaded:
-| ------------------------ | ---------- | ----- | --------------------------- |
+
-| Llama2                   | 7B         | 3.8GB | `ollama pull llama2`        |
+| Model              | Parameters | Size  | Download                       |
-| Llama2 13B               | 13B        | 7.3GB | `ollama pull llama2:13b`    |
+| ------------------ | ---------- | ----- | ------------------------------ |
-| Orca Mini                | 3B         | 1.9GB | `ollama pull orca`          |
+| Mistral            | 7B         | 4.1GB | `ollama run mistral`           |
-| Vicuna                   | 7B         | 3.8GB | `ollama pull vicuna`        |
+| Llama 2            | 7B         | 3.8GB | `ollama run llama2`            |
-| Nous-Hermes              | 13B        | 7.3GB | `ollama pull nous-hermes`   |
+| Code Llama         | 7B         | 3.8GB | `ollama run codellama`         |
-| Wizard Vicuna Uncensored | 13B        | 7.3GB | `ollama pull wizard-vicuna` |
+| Llama 2 Uncensored | 7B         | 3.8GB | `ollama run llama2-uncensored` |
 | Llama 2 13B        | 13B        | 7.3GB | `ollama run llama2:13b`        |
 | Llama 2 70B        | 70B        | 39GB  | `ollama run llama2:70b`        |
 | Orca Mini          | 3B         | 1.9GB | `ollama run orca-mini`         |
 | Vicuna             | 7B         | 3.8GB | `ollama run vicuna`            |
 > Note: You should have at least 8 GB of RAM to run the 3B models, 16 GB to run the 7B models, and 32 GB to run the 13B models.
-## Examples
+## Customize your own model
-### Run a model
+### Import from GGUF or GGML
-```
+Ollama supports importing GGUF and GGML file formats in the Modelfile. This means if you have a model that is not in the Ollama library, you can create it, iterate on it, and upload it to the Ollama library to share with others when you are ready.
 ollama run llama2
 >>> hi
 Hello! How can I help you today?
 ```
-### Create a custom model
+1. Create a file named Modelfile, and add a `FROM` instruction with the local filepath to the model you want to import.
-Pull a base model:
+   ```
   FROM ./vicuna-33b.Q4_0.gguf
   ```
 2. Create the model in Ollama
   ```
   ollama create name -f path_to_modelfile
   ```
 3. Run the model
   ```
   ollama run name
   ```
 ### Customize a prompt
 Models from the Ollama library can be customized with a prompt. The example
 ```
 ollama pull llama2
@@ -83,44 +109,85 @@ ollama run mario
 Hello! It's your friend Mario.
 ```
-For more examples, see the [examples](./examples) directory.
+For more examples, see the [examples](examples) directory. For more information on working with a Modelfile, see the [Modelfile](docs/modelfile.md) documentation.
-### Pull a model from the registry
+## CLI Reference
 ### Create a model
 `ollama create` is used to create a model from a Modelfile.
 ### Pull a model
 ```
-ollama pull orca
+ollama pull llama2
 ```
-### Listing local models
+> This command can also be used to update a local model. Only the diff will be pulled.
 ### Remove a model
 ```
 ollama rm llama2
 ```
 ### Copy a model
 ```
 ollama cp llama2 my-llama2
 ```
 ### Multiline input
 For multiline input, you can wrap text with `"""`:
 ```
 >>> """Hello,
 ... world!
 ... """
 I'm a basic program that prints the famous "Hello, world!" message to the console.
 ```
 ### Pass in prompt as arguments
 ```
 $ ollama run llama2 "summarize this file:" "$(cat README.md)"
 Ollama is a lightweight, extensible framework for building and running language models on the local machine. It provides a simple API for creating, running, and managing models, as well as a library of pre-built models that can be easily used in a variety of applications.
 ```
 ### List models on your computer
 ```
 ollama list
 ```
-## Model packages
+### Start Ollama
-### Overview
+`ollama serve` is used when you want to start ollama without running the desktop application.
 Ollama bundles model weights, configuration, and data into a single package, defined by a [Modelfile](./docs/modelfile.md).
 <picture>
  <source media="(prefers-color-scheme: dark)" height="480" srcset="https://github.com/jmorganca/ollama/assets/251292/2fd96b5f-191b-45c1-9668-941cfad4eb70">
  <img alt="logo" height="480" src="https://github.com/jmorganca/ollama/assets/251292/2fd96b5f-191b-45c1-9668-941cfad4eb70">
 </picture>
 ## Building
 Install `cmake` and `go`:
 ```
 brew install cmake
 brew install go
 ```
 Then generate dependencies and build:
 ```
 go generate ./...
 go build .
 ```
-To run it start the server:
+Next, start the server:
 ```
-./ollama serve &
+./ollama serve
 ```
-Finally, run a model!
+Finally, in a separate shell, run a model:
 ```
 ./ollama run llama2
@@ -128,10 +195,30 @@ Finally, run a model!
 ## REST API
-### `POST /api/generate`
+> See the [API documentation](docs/api.md) for all endpoints.
-Generate text from a model.
+Ollama has an API for running and managing models. For example to generate text from a model:
 ```
-curl -X POST http://localhost:11434/api/generate -d '{"model": "llama2", "prompt":"Why is the sky blue?"}'
+curl -X POST http://localhost:11434/api/generate -d '{
  "model": "llama2",
  "prompt":"Why is the sky blue?"
 }'
 ```
 ## Community Integrations
 - [LangChain](https://python.langchain.com/docs/integrations/llms/ollama) and [LangChain.js](https://js.langchain.com/docs/modules/model_io/models/llms/integrations/ollama) with [example](https://js.langchain.com/docs/use_cases/question_answering/local_retrieval_qa)
 - [LlamaIndex](https://gpt-index.readthedocs.io/en/stable/examples/llm/ollama.html)
 - [Raycast extension](https://github.com/MassimilianoPasquini97/raycast_ollama)
 - [Discollama](https://github.com/mxyng/discollama) (Discord bot inside the Ollama discord channel)
 - [Continue](https://github.com/continuedev/continue)
 - [Obsidian Ollama plugin](https://github.com/hinterdupfinger/obsidian-ollama)
 - [Dagger Chatbot](https://github.com/samalba/dagger-chatbot)
 - [LiteLLM](https://github.com/BerriAI/litellm)
 - [Discord AI Bot](https://github.com/mekb-turtle/discord-ai-bot)
 - [Chatbot UI](https://github.com/ivanfioravanti/chatbot-ollama)
 - [HTML UI](https://github.com/rtcfirefly/ollama-ui)
 - [Typescript UI](https://github.com/ollama-interface/Ollama-Gui?tab=readme-ov-file)
 - [Dumbar](https://github.com/JerrySievert/Dumbar)
 - [Emacs client](https://github.com/zweifisch/ollama)
--- a/api/client.go
+++ b/api/client.go
@@ -7,18 +7,27 @@ import (
 	"encoding/json"
 	"fmt"
 	"io"
 	"net"
 	"net/http"
 	"net/url"
 	"os"
 	"runtime"
 	"strings"
 	"github.com/jmorganca/ollama/version"
 )
 const DefaultHost = "127.0.0.1:11434"
 var envHost = os.Getenv("OLLAMA_HOST")
 type Client struct {
-	base    url.URL
+	base *url.URL
-	HTTP    http.Client
+	http http.Client
 	Headers http.Header
 }
 func checkError(resp *http.Response, body []byte) error {
-	if resp.StatusCode >= 200 && resp.StatusCode < 400 {
+	if resp.StatusCode < http.StatusBadRequest {
 		return nil
 	}
@@ -33,16 +42,44 @@ func checkError(resp *http.Response, body []byte) error {
 	return apiError
 }
-func NewClient(hosts ...string) *Client {
+func ClientFromEnvironment() (*Client, error) {
-	host := "127.0.0.1:11434"
+	scheme, hostport, ok := strings.Cut(os.Getenv("OLLAMA_HOST"), "://")
-	if len(hosts) > 0 {
+	if !ok {
-		host = hosts[0]
+		scheme, hostport = "http", os.Getenv("OLLAMA_HOST")
 	}
-	return &Client{
+	host, port, err := net.SplitHostPort(hostport)
-		base: url.URL{Scheme: "http", Host: host},
+	if err != nil {
-		HTTP: http.Client{},
+		host, port = "127.0.0.1", "11434"
 		if ip := net.ParseIP(strings.Trim(os.Getenv("OLLAMA_HOST"), "[]")); ip != nil {
 			host = ip.String()
 		}
 	}
 	client := Client{
 		base: &url.URL{
 			Scheme: scheme,
 			Host:   net.JoinHostPort(host, port),
 		},
 	}
 	mockRequest, err := http.NewRequest("HEAD", client.base.String(), nil)
 	if err != nil {
 		return nil, err
 	}
 	proxyURL, err := http.ProxyFromEnvironment(mockRequest)
 	if err != nil {
 		return nil, err
 	}
 	client.http = http.Client{
 		Transport: &http.Transport{
 			Proxy: http.ProxyURL(proxyURL),
 		},
 	}
 	return &client, nil
 }
 func (c *Client) do(ctx context.Context, method, path string, reqData, respData any) error {
@@ -57,21 +94,17 @@ func (c *Client) do(ctx context.Context, method, path string, reqData, respData
 		reqBody = bytes.NewReader(data)
 	}
-	url := c.base.JoinPath(path).String()
+	requestURL := c.base.JoinPath(path)
-
+	request, err := http.NewRequestWithContext(ctx, method, requestURL.String(), reqBody)
 	req, err := http.NewRequestWithContext(ctx, method, url, reqBody)
 	if err != nil {
 		return err
 	}
-	req.Header.Set("Content-Type", "application/json")
+	request.Header.Set("Content-Type", "application/json")
-	req.Header.Set("Accept", "application/json")
+	request.Header.Set("Accept", "application/json")
 	request.Header.Set("User-Agent", fmt.Sprintf("ollama/%s (%s %s) Go/%s", version.Version, runtime.GOARCH, runtime.GOOS, runtime.Version()))
-	for k, v := range c.Headers {
+	respObj, err := c.http.Do(request)
 		req.Header[k] = v
 	}
 	respObj, err := c.HTTP.Do(req)
 	if err != nil {
 		return err
 	}
@@ -94,6 +127,8 @@ func (c *Client) do(ctx context.Context, method, path string, reqData, respData
 	return nil
 }
 const maxBufferSize = 512 * 1000 // 512KB
 func (c *Client) stream(ctx context.Context, method, path string, data any, fn func([]byte) error) error {
 	var buf *bytes.Buffer
 	if data != nil {
@@ -105,21 +140,26 @@ func (c *Client) stream(ctx context.Context, method, path string, data any, fn f
 		buf = bytes.NewBuffer(bts)
 	}
-	request, err := http.NewRequestWithContext(ctx, method, c.base.JoinPath(path).String(), buf)
+	requestURL := c.base.JoinPath(path)
 	request, err := http.NewRequestWithContext(ctx, method, requestURL.String(), buf)
 	if err != nil {
 		return err
 	}
 	request.Header.Set("Content-Type", "application/json")
-	request.Header.Set("Accept", "application/json")
+	request.Header.Set("Accept", "application/x-ndjson")
 	request.Header.Set("User-Agent", fmt.Sprintf("ollama/%s (%s %s) Go/%s", version.Version, runtime.GOARCH, runtime.GOOS, runtime.Version()))
-	response, err := http.DefaultClient.Do(request)
+	response, err := c.http.Do(request)
 	if err != nil {
 		return err
 	}
 	defer response.Body.Close()
 	scanner := bufio.NewScanner(response.Body)
 	// increase the buffer size to avoid running out of space
 	scanBuf := make([]byte, 0, maxBufferSize)
 	scanner.Buffer(scanBuf, maxBufferSize)
 	for scanner.Scan() {
 		var errorResponse struct {
 			Error string `json:"error,omitempty"`
@@ -131,10 +171,10 @@ func (c *Client) stream(ctx context.Context, method, path string, data any, fn f
 		}
 		if errorResponse.Error != "" {
-			return fmt.Errorf("stream: %s", errorResponse.Error)
+			return fmt.Errorf(errorResponse.Error)
 		}
-		if response.StatusCode >= 400 {
+		if response.StatusCode >= http.StatusBadRequest {
 			return StatusError{
 				StatusCode:   response.StatusCode,
 				Status:       response.Status,
@@ -189,11 +229,11 @@ func (c *Client) Push(ctx context.Context, req *PushRequest, fn PushProgressFunc
 	})
 }
-type CreateProgressFunc func(CreateProgress) error
+type CreateProgressFunc func(ProgressResponse) error
 func (c *Client) Create(ctx context.Context, req *CreateRequest, fn CreateProgressFunc) error {
 	return c.stream(ctx, http.MethodPost, "/api/create", req, func(bts []byte) error {
-		var resp CreateProgress
+		var resp ProgressResponse
 		if err := json.Unmarshal(bts, &resp); err != nil {
 			return err
 		}
@@ -210,9 +250,31 @@ func (c *Client) List(ctx context.Context) (*ListResponse, error) {
 	return &lr, nil
 }
 func (c *Client) Copy(ctx context.Context, req *CopyRequest) error {
 	if err := c.do(ctx, http.MethodPost, "/api/copy", req, nil); err != nil {
 		return err
 	}
 	return nil
 }
 func (c *Client) Delete(ctx context.Context, req *DeleteRequest) error {
 	if err := c.do(ctx, http.MethodDelete, "/api/delete", req, nil); err != nil {
 		return err
 	}
 	return nil
 }
 func (c *Client) Show(ctx context.Context, req *ShowRequest) (*ShowResponse, error) {
 	var resp ShowResponse
 	if err := c.do(ctx, http.MethodPost, "/api/show", req, &resp); err != nil {
 		return nil, err
 	}
 	return &resp, nil
 }
 func (c *Client) Heartbeat(ctx context.Context) error {
 	if err := c.do(ctx, http.MethodHead, "/", nil, nil); err != nil {
 		return err
 	}
 	return nil
 }
--- a/api/client.py
+++ b/api/client.py
@@ -0,0 +1,225 @@
 import os
 import json
 import requests
 BASE_URL = os.environ.get('OLLAMA_HOST', 'http://localhost:11434')
 # Generate a response for a given prompt with a provided model. This is a streaming endpoint, so will be a series of responses.
 # The final response object will include statistics and additional data from the request. Use the callback function to override
 # the default handler.
 def generate(model_name, prompt, system=None, template=None, context=None, options=None, callback=None):
    try:
        url = f"{BASE_URL}/api/generate"
        payload = {
            "model": model_name, 
            "prompt": prompt, 
            "system": system, 
            "template": template, 
            "context": context, 
            "options": options
        }
        # Remove keys with None values
        payload = {k: v for k, v in payload.items() if v is not None}
        with requests.post(url, json=payload, stream=True) as response:
            response.raise_for_status()
            # Creating a variable to hold the context history of the final chunk
            final_context = None
            # Variable to hold concatenated response strings if no callback is provided
            full_response = ""
            # Iterating over the response line by line and displaying the details
            for line in response.iter_lines():
                if line:
                    # Parsing each line (JSON chunk) and extracting the details
                    chunk = json.loads(line)
                    # If a callback function is provided, call it with the chunk
                    if callback:
                        callback(chunk)
                    else:
                        # If this is not the last chunk, add the "response" field value to full_response and print it
                        if not chunk.get("done"):
                            response_piece = chunk.get("response", "")
                            full_response += response_piece
                            print(response_piece, end="", flush=True)
                    # Check if it's the last chunk (done is true)
                    if chunk.get("done"):
                        final_context = chunk.get("context")
            # Return the full response and the final context
            return full_response, final_context
    except requests.exceptions.RequestException as e:
        print(f"An error occurred: {e}")
        return None, None
 # Create a model from a Modelfile. Use the callback function to override the default handler.
 def create(model_name, model_path, callback=None):
    try:
        url = f"{BASE_URL}/api/create"
        payload = {"name": model_name, "path": model_path}
        # Making a POST request with the stream parameter set to True to handle streaming responses
        with requests.post(url, json=payload, stream=True) as response:
            response.raise_for_status()
            # Iterating over the response line by line and displaying the status
            for line in response.iter_lines():
                if line:
                    # Parsing each line (JSON chunk) and extracting the status
                    chunk = json.loads(line)
                    if callback:
                        callback(chunk)
                    else:
                        print(f"Status: {chunk.get('status')}")
    except requests.exceptions.RequestException as e:
        print(f"An error occurred: {e}")
 # Pull a model from a the model registry. Cancelled pulls are resumed from where they left off, and multiple
 # calls to will share the same download progress. Use the callback function to override the default handler.
 def pull(model_name, insecure=False, callback=None):
    try:
        url = f"{BASE_URL}/api/pull"
        payload = {
            "name": model_name,
            "insecure": insecure
        }
        # Making a POST request with the stream parameter set to True to handle streaming responses
        with requests.post(url, json=payload, stream=True) as response:
            response.raise_for_status()
            # Iterating over the response line by line and displaying the details
            for line in response.iter_lines():
                if line:
                    # Parsing each line (JSON chunk) and extracting the details
                    chunk = json.loads(line)
                    # If a callback function is provided, call it with the chunk
                    if callback:
                        callback(chunk)
                    else:
                        # Print the status message directly to the console
                        print(chunk.get('status', ''), end='', flush=True)
                    # If there's layer data, you might also want to print that (adjust as necessary)
                    if 'digest' in chunk:
                        print(f" - Digest: {chunk['digest']}", end='', flush=True)
                        print(f" - Total: {chunk['total']}", end='', flush=True)
                        print(f" - Completed: {chunk['completed']}", end='\n', flush=True)
                    else:
                        print()
    except requests.exceptions.RequestException as e:
        print(f"An error occurred: {e}")
 # Push a model to the model registry. Use the callback function to override the default handler.
 def push(model_name, insecure=False, callback=None):
    try:
        url = f"{BASE_URL}/api/push"
        payload = {
            "name": model_name,
            "insecure": insecure
        }
        # Making a POST request with the stream parameter set to True to handle streaming responses
        with requests.post(url, json=payload, stream=True) as response:
            response.raise_for_status()
            # Iterating over the response line by line and displaying the details
            for line in response.iter_lines():
                if line:
                    # Parsing each line (JSON chunk) and extracting the details
                    chunk = json.loads(line)
                    # If a callback function is provided, call it with the chunk
                    if callback:
                        callback(chunk)
                    else:
                        # Print the status message directly to the console
                        print(chunk.get('status', ''), end='', flush=True)
                    # If there's layer data, you might also want to print that (adjust as necessary)
                    if 'digest' in chunk:
                        print(f" - Digest: {chunk['digest']}", end='', flush=True)
                        print(f" - Total: {chunk['total']}", end='', flush=True)
                        print(f" - Completed: {chunk['completed']}", end='\n', flush=True)
                    else:
                        print()
    except requests.exceptions.RequestException as e:
        print(f"An error occurred: {e}")
 # List models that are available locally.
 def list():
    try:
        response = requests.get(f"{BASE_URL}/api/tags")
        response.raise_for_status()
        data = response.json()
        models = data.get('models', [])
        return models
    except requests.exceptions.RequestException as e:
        print(f"An error occurred: {e}")
        return None
 # Copy a model. Creates a model with another name from an existing model.
 def copy(source, destination):
    try:
        # Create the JSON payload
        payload = {
            "source": source,
            "destination": destination
        }
        response = requests.post(f"{BASE_URL}/api/copy", json=payload)
        response.raise_for_status()
        # If the request was successful, return a message indicating that the copy was successful
        return "Copy successful"
    except requests.exceptions.RequestException as e:
        print(f"An error occurred: {e}")
        return None
 # Delete a model and its data.
 def delete(model_name):
    try:
        url = f"{BASE_URL}/api/delete"
        payload = {"name": model_name}
        response = requests.delete(url, json=payload)
        response.raise_for_status()
        return "Delete successful"
    except requests.exceptions.RequestException as e:
        print(f"An error occurred: {e}")
        return None
 # Show info about a model.
 def show(model_name):
    try:
        url = f"{BASE_URL}/api/show"
        payload = {"name": model_name}
        response = requests.post(url, json=payload)
        response.raise_for_status()
        # Parse the JSON response and return it
        data = response.json()
        return data
    except requests.exceptions.RequestException as e:
        print(f"An error occurred: {e}")
        return None
 def heartbeat():
    try:
        url = f"{BASE_URL}/"
        response = requests.head(url)
        response.raise_for_status()
        return "Ollama is running"
    except requests.exceptions.RequestException as e:
        print(f"An error occurred: {e}")
        return "Ollama is not running"
--- a/api/types.go
+++ b/api/types.go
@@ -1,9 +1,13 @@
 package api
 import (
 	"encoding/json"
 	"fmt"
 	"log"
 	"math"
 	"os"
-	"runtime"
+	"reflect"
 	"strings"
 	"time"
 )
@@ -28,38 +32,67 @@ func (e StatusError) Error() string {
 }
 type GenerateRequest struct {
-	Model   string `json:"model"`
+	Model    string `json:"model"`
-	Prompt  string `json:"prompt"`
+	Prompt   string `json:"prompt"`
-	Context []int  `json:"context,omitempty"`
+	System   string `json:"system"`
 	Template string `json:"template"`
 	Context  []int  `json:"context,omitempty"`
 	Stream   *bool  `json:"stream,omitempty"`
-	Options `json:"options"`
+	Options map[string]interface{} `json:"options"`
 }
 type EmbeddingRequest struct {
 	Model  string `json:"model"`
 	Prompt string `json:"prompt"`
 	Options map[string]interface{} `json:"options"`
 }
 type EmbeddingResponse struct {
 	Embedding []float64 `json:"embedding"`
 }
 type CreateRequest struct {
-	Name string `json:"name"`
+	Name   string `json:"name"`
-	Path string `json:"path"`
+	Path   string `json:"path"`
-}
+	Stream *bool  `json:"stream,omitempty"`
 type CreateProgress struct {
 	Status string `json:"status"`
 }
 type DeleteRequest struct {
 	Name string `json:"name"`
 }
 type ShowRequest struct {
 	Name string `json:"name"`
 }
 type ShowResponse struct {
 	License    string `json:"license,omitempty"`
 	Modelfile  string `json:"modelfile,omitempty"`
 	Parameters string `json:"parameters,omitempty"`
 	Template   string `json:"template,omitempty"`
 	System     string `json:"system,omitempty"`
 }
 type CopyRequest struct {
 	Source      string `json:"source"`
 	Destination string `json:"destination"`
 }
 type PullRequest struct {
 	Name     string `json:"name"`
 	Insecure bool   `json:"insecure,omitempty"`
 	Username string `json:"username"`
 	Password string `json:"password"`
 	Stream   *bool  `json:"stream,omitempty"`
 }
 type ProgressResponse struct {
 	Status    string `json:"status"`
 	Digest    string `json:"digest,omitempty"`
-	Total     int    `json:"total,omitempty"`
+	Total     int64  `json:"total,omitempty"`
-	Completed int    `json:"completed,omitempty"`
+	Completed int64  `json:"completed,omitempty"`
 }
 type PushRequest struct {
@@ -67,27 +100,34 @@ type PushRequest struct {
 	Insecure bool   `json:"insecure,omitempty"`
 	Username string `json:"username"`
 	Password string `json:"password"`
 	Stream   *bool  `json:"stream,omitempty"`
 }
 type ListResponse struct {
-	Models []ListResponseModel `json:"models"`
+	Models []ModelResponse `json:"models"`
 }
-type ListResponseModel struct {
+type ModelResponse struct {
 	Name       string    `json:"name"`
 	ModifiedAt time.Time `json:"modified_at"`
-	Size       int       `json:"size"`
+	Size       int64     `json:"size"`
 	Digest     string    `json:"digest"`
 }
 type TokenResponse struct {
 	Token string `json:"token"`
 }
 type GenerateResponse struct {
 	Model     string    `json:"model"`
 	CreatedAt time.Time `json:"created_at"`
-	Response  string    `json:"response,omitempty"`
+	Response  string    `json:"response"`
 	Done    bool  `json:"done"`
 	Context []int `json:"context,omitempty"`
 	TotalDuration      time.Duration `json:"total_duration,omitempty"`
 	LoadDuration       time.Duration `json:"load_duration,omitempty"`
 	PromptEvalCount    int           `json:"prompt_eval_count,omitempty"`
 	PromptEvalDuration time.Duration `json:"prompt_eval_duration,omitempty"`
 	EvalCount          int           `json:"eval_count,omitempty"`
@@ -99,6 +139,10 @@ func (r *GenerateResponse) Summary() {
 		fmt.Fprintf(os.Stderr, "total duration:       %v\n", r.TotalDuration)
 	}
 	if r.LoadDuration > 0 {
 		fmt.Fprintf(os.Stderr, "load duration:        %v\n", r.LoadDuration)
 	}
 	if r.PromptEvalCount > 0 {
 		fmt.Fprintf(os.Stderr, "prompt eval count:    %d token(s)\n", r.PromptEvalCount)
 	}
@@ -125,62 +169,194 @@ type Options struct {
 	UseNUMA bool `json:"numa,omitempty"`
 	// Model options
-	NumCtx        int  `json:"num_ctx,omitempty"`
+	NumCtx             int     `json:"num_ctx,omitempty"`
-	NumBatch      int  `json:"num_batch,omitempty"`
+	NumKeep            int     `json:"num_keep,omitempty"`
-	NumGPU        int  `json:"num_gpu,omitempty"`
+	NumBatch           int     `json:"num_batch,omitempty"`
-	MainGPU       int  `json:"main_gpu,omitempty"`
+	NumGQA             int     `json:"num_gqa,omitempty"`
-	LowVRAM       bool `json:"low_vram,omitempty"`
+	NumGPU             int     `json:"num_gpu,omitempty"`
-	F16KV         bool `json:"f16_kv,omitempty"`
+	MainGPU            int     `json:"main_gpu,omitempty"`
-	LogitsAll     bool `json:"logits_all,omitempty"`
+	LowVRAM            bool    `json:"low_vram,omitempty"`
-	VocabOnly     bool `json:"vocab_only,omitempty"`
+	F16KV              bool    `json:"f16_kv,omitempty"`
-	UseMMap       bool `json:"use_mmap,omitempty"`
+	LogitsAll          bool    `json:"logits_all,omitempty"`
-	UseMLock      bool `json:"use_mlock,omitempty"`
+	VocabOnly          bool    `json:"vocab_only,omitempty"`
-	EmbeddingOnly bool `json:"embedding_only,omitempty"`
+	UseMMap            bool    `json:"use_mmap,omitempty"`
 	UseMLock           bool    `json:"use_mlock,omitempty"`
 	EmbeddingOnly      bool    `json:"embedding_only,omitempty"`
 	RopeFrequencyBase  float32 `json:"rope_frequency_base,omitempty"`
 	RopeFrequencyScale float32 `json:"rope_frequency_scale,omitempty"`
 	// Predict options
-	RepeatLastN      int     `json:"repeat_last_n,omitempty"`
+	NumPredict       int      `json:"num_predict,omitempty"`
-	RepeatPenalty    float32 `json:"repeat_penalty,omitempty"`
+	TopK             int      `json:"top_k,omitempty"`
-	FrequencyPenalty float32 `json:"frequency_penalty,omitempty"`
+	TopP             float32  `json:"top_p,omitempty"`
-	PresencePenalty  float32 `json:"presence_penalty,omitempty"`
+	TFSZ             float32  `json:"tfs_z,omitempty"`
-	Temperature      float32 `json:"temperature,omitempty"`
+	TypicalP         float32  `json:"typical_p,omitempty"`
-	TopK             int     `json:"top_k,omitempty"`
+	RepeatLastN      int      `json:"repeat_last_n,omitempty"`
-	TopP             float32 `json:"top_p,omitempty"`
+	Temperature      float32  `json:"temperature,omitempty"`
-	TFSZ             float32 `json:"tfs_z,omitempty"`
+	RepeatPenalty    float32  `json:"repeat_penalty,omitempty"`
-	TypicalP         float32 `json:"typical_p,omitempty"`
+	PresencePenalty  float32  `json:"presence_penalty,omitempty"`
-	Mirostat         int     `json:"mirostat,omitempty"`
+	FrequencyPenalty float32  `json:"frequency_penalty,omitempty"`
-	MirostatTau      float32 `json:"mirostat_tau,omitempty"`
+	Mirostat         int      `json:"mirostat,omitempty"`
-	MirostatEta      float32 `json:"mirostat_eta,omitempty"`
+	MirostatTau      float32  `json:"mirostat_tau,omitempty"`
 	MirostatEta      float32  `json:"mirostat_eta,omitempty"`
 	PenalizeNewline  bool     `json:"penalize_newline,omitempty"`
 	Stop             []string `json:"stop,omitempty"`
 	NumThread int `json:"num_thread,omitempty"`
 }
 var ErrInvalidOpts = fmt.Errorf("invalid options")
 func (opts *Options) FromMap(m map[string]interface{}) error {
 	valueOpts := reflect.ValueOf(opts).Elem() // names of the fields in the options struct
 	typeOpts := reflect.TypeOf(opts).Elem()   // types of the fields in the options struct
 	// build map of json struct tags to their types
 	jsonOpts := make(map[string]reflect.StructField)
 	for _, field := range reflect.VisibleFields(typeOpts) {
 		jsonTag := strings.Split(field.Tag.Get("json"), ",")[0]
 		if jsonTag != "" {
 			jsonOpts[jsonTag] = field
 		}
 	}
 	invalidOpts := []string{}
 	for key, val := range m {
 		if opt, ok := jsonOpts[key]; ok {
 			field := valueOpts.FieldByName(opt.Name)
 			if field.IsValid() && field.CanSet() {
 				if val == nil {
 					continue
 				}
 				switch field.Kind() {
 				case reflect.Int:
 					switch t := val.(type) {
 					case int64:
 						field.SetInt(t)
 					case float64:
 						// when JSON unmarshals numbers, it uses float64, not int
 						field.SetInt(int64(t))
 					default:
 						log.Printf("could not convert model parameter %v of type %T to int, skipped", key, val)
 					}
 				case reflect.Bool:
 					val, ok := val.(bool)
 					if !ok {
 						log.Printf("could not convert model parameter %v of type %T to bool, skipped", key, val)
 						continue
 					}
 					field.SetBool(val)
 				case reflect.Float32:
 					// JSON unmarshals to float64
 					val, ok := val.(float64)
 					if !ok {
 						log.Printf("could not convert model parameter %v of type %T to float32, skipped", key, val)
 						continue
 					}
 					field.SetFloat(val)
 				case reflect.String:
 					val, ok := val.(string)
 					if !ok {
 						log.Printf("could not convert model parameter %v of type %T to string, skipped", key, val)
 						continue
 					}
 					field.SetString(val)
 				case reflect.Slice:
 					// JSON unmarshals to []interface{}, not []string
 					val, ok := val.([]interface{})
 					if !ok {
 						log.Printf("could not convert model parameter %v of type %T to slice, skipped", key, val)
 						continue
 					}
 					// convert []interface{} to []string
 					slice := make([]string, len(val))
 					for i, item := range val {
 						str, ok := item.(string)
 						if !ok {
 							log.Printf("could not convert model parameter %v of type %T to slice of strings, skipped", key, item)
 							continue
 						}
 						slice[i] = str
 					}
 					field.Set(reflect.ValueOf(slice))
 				default:
 					return fmt.Errorf("unknown type loading config params: %v", field.Kind())
 				}
 			}
 		} else {
 			invalidOpts = append(invalidOpts, key)
 		}
 	}
 	if len(invalidOpts) > 0 {
 		return fmt.Errorf("%w: %v", ErrInvalidOpts, strings.Join(invalidOpts, ", "))
 	}
 	return nil
 }
 func DefaultOptions() Options {
 	return Options{
-		Seed: -1,
+		// options set on request to runner
-
+		NumPredict:       -1,
-		UseNUMA: false,
+		NumKeep:          -1,
 		NumCtx:   2048,
 		NumBatch: 512,
 		NumGPU:   1,
 		LowVRAM:  false,
 		F16KV:    true,
 		UseMMap:  true,
 		UseMLock: false,
 		RepeatLastN:      512,
 		RepeatPenalty:    1.1,
 		FrequencyPenalty: 0.0,
 		PresencePenalty:  0.0,
 		Temperature:      0.8,
 		TopK:             40,
 		TopP:             0.9,
 		TFSZ:             1.0,
 		TypicalP:         1.0,
 		RepeatLastN:      64,
 		RepeatPenalty:    1.1,
 		PresencePenalty:  0.0,
 		FrequencyPenalty: 0.0,
 		Mirostat:         0,
 		MirostatTau:      5.0,
 		MirostatEta:      0.1,
 		PenalizeNewline:  true,
 		Seed:             -1,
-		NumThread: runtime.NumCPU(),
+		// options set when the model is loaded
 		NumCtx:             2048,
 		RopeFrequencyBase:  10000.0,
 		RopeFrequencyScale: 1.0,
 		NumBatch:           512,
 		NumGPU:             -1, // -1 here indicates that NumGPU should be set dynamically
 		NumGQA:             1,
 		NumThread:          0, // let the runtime decide
 		LowVRAM:            false,
 		F16KV:              true,
 		UseMLock:           false,
 		UseMMap:            true,
 		UseNUMA:            false,
 		EmbeddingOnly:      true,
 	}
 }
 type Duration struct {
 	time.Duration
 }
 func (d *Duration) UnmarshalJSON(b []byte) (err error) {
 	var v any
 	if err := json.Unmarshal(b, &v); err != nil {
 		return err
 	}
 	d.Duration = 5 * time.Minute
 	switch t := v.(type) {
 	case float64:
 		if t < 0 {
 			t = math.MaxFloat64
 		}
 		d.Duration = time.Duration(t)
 	case string:
 		d.Duration, err = time.ParseDuration(t)
 		if err != nil {
 			return err
 		}
 	}
 	return nil
 }
--- a/app/README.md
+++ b/app/README.md
@@ -1,7 +1,5 @@
 # Desktop
 _Note: the Ollama desktop app is a work in progress and is not ready yet for general use._
 This app builds upon Ollama to provide a desktop experience for running models.
 ## Developing
@@ -9,19 +7,15 @@ This app builds upon Ollama to provide a desktop experience for running models.
 First, build the `ollama` binary:
 ```
-make -C ..
+cd ..
 go build .
 ```
 Then run the desktop app with `npm start`:
 ```
 cd app
 npm install
 npm start
 ```
 ## Coming soon
 - Browse the latest available models on Hugging Face and other sources
 - Keep track of previous conversations with models
 - Switch quickly between models
 - Connect to remote Ollama servers to run models
--- a/app/assets/iconDarkTemplate.png
+++ b/app/assets/iconDarkTemplate.png
--- a/app/assets/ollama_icon_16x16Template@2x.png
+++ b/app/assets/ollama_icon_16x16Template@2x.png
--- a/app/assets/iconDarkUpdateTemplate.png
+++ b/app/assets/iconDarkUpdateTemplate.png
--- a/app/assets/iconDarkUpdateTemplate@2x.png
+++ b/app/assets/iconDarkUpdateTemplate@2x.png
--- a/app/assets/iconTemplate.png
+++ b/app/assets/iconTemplate.png
--- a/app/assets/ollama_outline_icon_16x16Template@2x.png
+++ b/app/assets/ollama_outline_icon_16x16Template@2x.png
--- a/app/assets/iconUpdateTemplate.png
+++ b/app/assets/iconUpdateTemplate.png
--- a/app/assets/iconUpdateTemplate@2x.png
+++ b/app/assets/iconUpdateTemplate@2x.png
--- a/app/assets/ollama_icon_16x16Template.png
+++ b/app/assets/ollama_icon_16x16Template.png
--- a/app/assets/ollama_outline_icon_16x16Template.png
+++ b/app/assets/ollama_outline_icon_16x16Template.png
--- a/app/forge.config.ts
+++ b/app/forge.config.ts
@@ -18,12 +18,15 @@ const config: ForgeConfig = {
    asar: true,
    icon: './assets/icon.icns',
    extraResource: [
-      '../ollama',
+      '../dist/ollama',
-      path.join(__dirname, './assets/ollama_icon_16x16Template.png'),
+      path.join(__dirname, './assets/iconTemplate.png'),
-      path.join(__dirname, './assets/ollama_icon_16x16Template@2x.png'),
+      path.join(__dirname, './assets/iconTemplate@2x.png'),
-      path.join(__dirname, './assets/ollama_outline_icon_16x16Template.png'),
+      path.join(__dirname, './assets/iconUpdateTemplate.png'),
-      path.join(__dirname, './assets/ollama_outline_icon_16x16Template@2x.png'),
+      path.join(__dirname, './assets/iconUpdateTemplate@2x.png'),
-      ...(process.platform === 'darwin' ? ['../llama/ggml-metal.metal'] : []),
+      path.join(__dirname, './assets/iconDarkTemplate.png'),
      path.join(__dirname, './assets/iconDarkTemplate@2x.png'),
      path.join(__dirname, './assets/iconDarkUpdateTemplate.png'),
      path.join(__dirname, './assets/iconDarkUpdateTemplate@2x.png'),
    ],
    ...(process.env.SIGN
      ? {
@@ -38,6 +41,9 @@ const config: ForgeConfig = {
          },
        }
      : {}),
    osxUniversal: {
      x64ArchFiles: '**/ollama',
    },
  },
  rebuildConfig: {},
  makers: [new MakerSquirrel({}), new MakerZIP({}, ['darwin'])],
--- a/app/package-lock.json
+++ b/app/package-lock.json
@@ -32,6 +32,7 @@
        "@electron-forge/plugin-auto-unpack-natives": "^6.2.1",
        "@electron-forge/plugin-webpack": "^6.2.1",
        "@electron-forge/publisher-github": "^6.2.1",
        "@electron/universal": "^1.4.1",
        "@svgr/webpack": "^8.0.1",
        "@types/chmodr": "^1.0.0",
        "@types/node": "^20.4.0",
@@ -3328,9 +3329,9 @@
      }
    },
    "node_modules/@electron/universal": {
-      "version": "1.3.4",
+      "version": "1.4.1",
-      "resolved": "https://registry.npmjs.org/@electron/universal/-/universal-1.3.4.tgz",
+      "resolved": "https://registry.npmjs.org/@electron/universal/-/universal-1.4.1.tgz",
-      "integrity": "sha512-BdhBgm2ZBnYyYRLRgOjM5VHkyFItsbggJ0MHycOjKWdFGYwK97ZFXH54dTvUWEfha81vfvwr5On6XBjt99uDcg==",
+      "integrity": "sha512-lE/U3UNw1YHuowNbTmKNs9UlS3En3cPgwM5MI+agIgr/B1hSze9NdOP0qn7boZaI9Lph8IDv3/24g9IxnJP7aQ==",
      "dev": true,
      "dependencies": {
        "@electron/asar": "^3.2.1",
--- a/app/package.json
+++ b/app/package.json
@@ -6,10 +6,10 @@
  "main": ".webpack/main",
  "scripts": {
    "start": "electron-forge start",
-    "package": "electron-forge package",
+    "package": "electron-forge package --arch universal",
-    "package:sign": "SIGN=1 electron-forge package",
+    "package:sign": "SIGN=1 electron-forge package --arch universal",
-    "make": "electron-forge make",
+    "make": "electron-forge make --arch universal",
-    "make:sign": "SIGN=1 electron-forge make",
+    "make:sign": "SIGN=1 electron-forge make --arch universal",
    "publish": "SIGN=1 electron-forge publish",
    "lint": "eslint --ext .ts,.tsx .",
    "format": "prettier --check . --ignore-path .gitignore",
@@ -32,6 +32,7 @@
    "@electron-forge/plugin-auto-unpack-natives": "^6.2.1",
    "@electron-forge/plugin-webpack": "^6.2.1",
    "@electron-forge/publisher-github": "^6.2.1",
    "@electron/universal": "^1.4.1",
    "@svgr/webpack": "^8.0.1",
    "@types/chmodr": "^1.0.0",
    "@types/node": "^20.4.0",
--- a/app/src/app.tsx
+++ b/app/src/app.tsx
@@ -2,7 +2,7 @@ import { useState } from 'react'
 import copy from 'copy-to-clipboard'
 import { CheckIcon, DocumentDuplicateIcon } from '@heroicons/react/24/outline'
 import Store from 'electron-store'
-import { getCurrentWindow } from '@electron/remote'
+import { getCurrentWindow, app } from '@electron/remote'
 import { install } from './install'
 import OllamaIcon from './ollama.svg'
@@ -51,10 +51,15 @@ export default function () {
              <div className='mx-auto'>
                <button
                  onClick={async () => {
-                    await install()
+                    try {
-                    getCurrentWindow().show()
+                      await install()
-                    getCurrentWindow().focus()
+                      setStep(Step.FINISH)
-                    setStep(Step.FINISH)
+                    } catch (e) {
                      console.error('could not install: ', e)
                    } finally {
                      getCurrentWindow().show()
                      getCurrentWindow().focus()
                    }
                  }}
                  className='no-drag rounded-dm mx-auto w-[60%] rounded-md bg-black px-4 py-2 text-sm text-white hover:brightness-110'
                >
--- a/app/src/index.ts
+++ b/app/src/index.ts
@@ -1,17 +1,21 @@
-import { spawn } from 'child_process'
+import { spawn, ChildProcess } from 'child_process'
-import { app, autoUpdater, dialog, Tray, Menu, BrowserWindow, nativeTheme } from 'electron'
+import { app, autoUpdater, dialog, Tray, Menu, BrowserWindow, MenuItemConstructorOptions, nativeTheme } from 'electron'
 import Store from 'electron-store'
 import winston from 'winston'
 import 'winston-daily-rotate-file'
 import * as path from 'path'
-import { analytics, id } from './telemetry'
+import { v4 as uuidv4 } from 'uuid'
 import { installed } from './install'
 require('@electron/remote/main').initialize()
 if (require('electron-squirrel-startup')) {
  app.quit()
 }
 const store = new Store()
-let tray: Tray | null = null
+
 let welcomeWindow: BrowserWindow | null = null
 declare const MAIN_WINDOW_WEBPACK_ENTRY: string
@@ -28,10 +32,30 @@ const logger = winston.createLogger({
  format: winston.format.printf(info => info.message),
 })
-const SingleInstanceLock = app.requestSingleInstanceLock()
+app.on('ready', () => {
-if (!SingleInstanceLock) {
+  const gotTheLock = app.requestSingleInstanceLock()
-  app.quit()
+  if (!gotTheLock) {
-}
+    app.exit(0)
    return
  }
  app.on('second-instance', () => {
    if (app.hasSingleInstanceLock()) {
      app.releaseSingleInstanceLock()
    }
    if (proc) {
      proc.off('exit', restart)
      proc.kill()
    }
    app.exit(0)
  })
  app.focus({ steal: true })
  init()
 })
 function firstRunWindow() {
  // Create the browser window.
@@ -47,65 +71,74 @@ function firstRunWindow() {
      nodeIntegration: true,
      contextIsolation: false,
    },
    alwaysOnTop: true,
  })
  require('@electron/remote/main').enable(welcomeWindow.webContents)
  // and load the index.html of the app.
  welcomeWindow.loadURL(MAIN_WINDOW_WEBPACK_ENTRY)
  welcomeWindow.on('ready-to-show', () => welcomeWindow.show())
-
+  welcomeWindow.on('closed', () => {
-  // for debugging
+    if (process.platform === 'darwin') {
-  // welcomeWindow.webContents.openDevTools()
+      app.dock.hide()
  if (process.platform === 'darwin') {
    app.dock.hide()
  }
 }
 function createSystemtray() {
  let iconPath = nativeTheme.shouldUseDarkColors
    ? path.join(__dirname, '..', '..', 'assets', 'ollama_icon_16x16Template.png')
    : path.join(__dirname, '..', '..', 'assets', 'ollama_outline_icon_16x16Template.png')
  if (app.isPackaged) {
    iconPath = nativeTheme.shouldUseDarkColors
      ? path.join(process.resourcesPath, 'ollama_icon_16x16Template.png')
      : path.join(process.resourcesPath, 'ollama_outline_icon_16x16Template.png')
  }
  tray = new Tray(iconPath)
  nativeTheme.on('updated', function theThemeHasChanged() {
    if (nativeTheme.shouldUseDarkColors) {
      app.isPackaged
        ? tray.setImage(path.join(process.resourcesPath, 'ollama_icon_16x16Template.png'))
        : tray.setImage(path.join(__dirname, '..', '..', 'assets', 'ollama_icon_16x16Template.png'))
    } else {
      app.isPackaged
        ? tray.setImage(path.join(process.resourcesPath, 'ollama_outline_icon_16x16Template.png'))
        : tray.setImage(path.join(__dirname, '..', '..', 'assets', 'ollama_outline_icon_16x16Template.png'))
    }
  })
  const contextMenu = Menu.buildFromTemplate([{ role: 'quit', label: 'Quit Ollama', accelerator: 'Command+Q' }])
  tray.setContextMenu(contextMenu)
  tray.setToolTip('Ollama')
 }
-if (require('electron-squirrel-startup')) {
+let tray: Tray | null = null
-  app.quit()
+let updateAvailable = false
 const assetPath = app.isPackaged ? process.resourcesPath : path.join(__dirname, '..', '..', 'assets')
 function trayIconPath() {
  return nativeTheme.shouldUseDarkColors
    ? updateAvailable
      ? path.join(assetPath, 'iconDarkUpdateTemplate.png')
      : path.join(assetPath, 'iconDarkTemplate.png')
    : updateAvailable
    ? path.join(assetPath, 'iconUpdateTemplate.png')
    : path.join(assetPath, 'iconTemplate.png')
 }
 function updateTrayIcon() {
  if (tray) {
    tray.setImage(trayIconPath())
  }
 }
 function updateTray() {
  const updateItems: MenuItemConstructorOptions[] = [
    { label: 'An update is available', enabled: false },
    {
      label: 'Restart to update',
      click: () => autoUpdater.quitAndInstall(),
    },
    { type: 'separator' },
  ]
  const menu = Menu.buildFromTemplate([
    ...(updateAvailable ? updateItems : []),
    { role: 'quit', label: 'Quit Ollama', accelerator: 'Command+Q' },
  ])
  if (!tray) {
    tray = new Tray(trayIconPath())
  }
  tray.setToolTip(updateAvailable ? 'An update is available' : 'Ollama')
  tray.setContextMenu(menu)
  tray.setImage(trayIconPath())
  nativeTheme.off('updated', updateTrayIcon)
  nativeTheme.on('updated', updateTrayIcon)
 }
 let proc: ChildProcess = null
 function server() {
  const binary = app.isPackaged
    ? path.join(process.resourcesPath, 'ollama')
    : path.resolve(process.cwd(), '..', 'ollama')
-  const proc = spawn(binary, ['serve'])
+  proc = spawn(binary, ['serve'])
  proc.stdout.on('data', data => {
    logger.info(data.toString().trim())
@@ -115,23 +148,32 @@ function server() {
    logger.error(data.toString().trim())
  })
-  function restart() {
+  proc.on('exit', restart)
-    setTimeout(server, 3000)
+}
 function restart() {
  setTimeout(server, 1000)
 }
 app.on('before-quit', () => {
  if (proc) {
    proc.off('exit', restart)
    proc.kill('SIGINT') // send SIGINT signal to the server, which also stops any loaded llms
  }
 })
 function init() {
  if (app.isPackaged) {
    autoUpdater.checkForUpdates()
    setInterval(() => {
      if (!updateAvailable) {
        autoUpdater.checkForUpdates()
      }
    }, 60 * 60 * 1000)
  }
-  proc.on('exit', restart)
+  updateTray()
  app.on('before-quit', () => {
    proc.off('exit', restart)
    proc.kill()
  })
 }
 if (process.platform === 'darwin') {
  app.dock.hide()
 }
 app.on('ready', () => {
  if (process.platform === 'darwin') {
    if (app.isPackaged) {
      if (!app.isInApplicationsFolder()) {
@@ -167,10 +209,13 @@ app.on('ready', () => {
    }
  }
  createSystemtray()
  server()
  if (store.get('first-time-run') && installed()) {
    if (process.platform === 'darwin') {
      app.dock.hide()
    }
    app.setLoginItemSettings({ openAtLogin: app.getLoginItemSettings().openAtLogin })
    return
  }
@@ -178,7 +223,7 @@ app.on('ready', () => {
  // This is the first run or the CLI is no longer installed
  app.setLoginItemSettings({ openAtLogin: true })
  firstRunWindow()
-})
+}
 // Quit when all windows are closed, except on macOS. There, it's common
 // for applications and their menu bar to stay active until the user quits
@@ -189,45 +234,30 @@ app.on('window-all-closed', () => {
  }
 })
-// In this file you can include the rest of your app's specific main process
+function id(): string {
-// code. You can also put them in separate files and import them here.
+  const id = store.get('id') as string
  if (id) {
    return id
  }
  const uuid = uuidv4()
  store.set('id', uuid)
  return uuid
 }
 autoUpdater.setFeedURL({
-  url: `https://ollama.ai/api/update?os=${process.platform}&arch=${process.arch}&version=${app.getVersion()}`,
+  url: `https://ollama.ai/api/update?os=${process.platform}&arch=${
    process.arch
  }&version=${app.getVersion()}&id=${id()}`,
 })
 async function heartbeat() {
  analytics.track({
    anonymousId: id(),
    event: 'heartbeat',
    properties: {
      version: app.getVersion(),
    },
  })
 }
 if (app.isPackaged) {
  heartbeat()
  autoUpdater.checkForUpdates()
  setInterval(() => {
    heartbeat()
    autoUpdater.checkForUpdates()
  }, 60 * 60 * 1000)
 }
 autoUpdater.on('error', e => {
  logger.error(`update check failed - ${e.message}`)
  console.error(`update check failed - ${e.message}`)
 })
-autoUpdater.on('update-downloaded', (event, releaseNotes, releaseName) => {
+autoUpdater.on('update-downloaded', () => {
-  dialog
+  updateAvailable = true
-    .showMessageBox({
+  updateTray()
      type: 'info',
      buttons: ['Restart Now', 'Later'],
      title: 'New update available',
      message: process.platform === 'win32' ? releaseNotes : releaseName,
      detail: 'A new version of Ollama is available. Restart to apply the update.',
    })
    .then(returnValue => {
      if (returnValue.response === 0) autoUpdater.quitAndInstall()
    })
 })
--- a/app/src/install.ts
+++ b/app/src/install.ts
@@ -15,12 +15,7 @@ export function installed() {
 export async function install() {
  const command = `do shell script "mkdir -p ${path.dirname(
    symlinkPath
-  )} && ln -F -s ${ollama} ${symlinkPath}" with administrator privileges`
+  )} && ln -F -s \\"${ollama}\\" \\"${symlinkPath}\\"" with administrator privileges`
-  try {
+  await exec(`osascript -e '${command}'`)
    await exec(`osascript -e '${command}'`)
  } catch (error) {
    console.error(`cli: failed to install cli: ${error.message}`)
    return
  }
 }
--- a/app/src/telemetry.ts
+++ b/app/src/telemetry.ts
@@ -1,19 +0,0 @@
 import { Analytics } from '@segment/analytics-node'
 import { v4 as uuidv4 } from 'uuid'
 import Store from 'electron-store'
 const store = new Store()
 export const analytics = new Analytics({ writeKey: process.env.TELEMETRY_WRITE_KEY || '<empty>' })
 export function id(): string {
  const id = store.get('id') as string
  if (id) {
    return id
  }
  const uuid = uuidv4()
  store.set('id', uuid)
  return uuid
 }
--- a/cmd/cmd.go
+++ b/cmd/cmd.go
--- a/docs/README.md
+++ b/docs/README.md
@@ -0,0 +1,6 @@
 # Documentation
 - [Modelfile](./modelfile.md)
 - [How to develop Ollama](./development.md)
 - [API](./api.md)
 - [Tutorials](./tutorials.md)
--- a/docs/api.md
+++ b/docs/api.md
@@ -0,0 +1,363 @@
 # API
 ## Endpoints
 - [Generate a completion](#generate-a-completion)
 - [Create a Model](#create-a-model)
 - [List Local Models](#list-local-models)
 - [Show Model Information](#show-model-information)
 - [Copy a Model](#copy-a-model)
 - [Delete a Model](#delete-a-model)
 - [Pull a Model](#pull-a-model)
 - [Push a Model](#push-a-model)
 - [Generate Embeddings](#generate-embeddings)
 ## Conventions
 ### Model names
 Model names follow a `model:tag` format. Some examples are `orca-mini:3b-q4_1` and `llama2:70b`. The tag is optional and, if not provided, will default to `latest`. The tag is used to identify a specific version.
 ### Durations
 All durations are returned in nanoseconds.
 ### Streaming responses
 Certain endpoints stream responses as JSON objects delineated with the newline (`\n`) character.
 ## Generate a completion
 ```shell
 POST /api/generate
 ```
 Generate a response for a given prompt with a provided model. This is a streaming endpoint, so will be a series of responses. The final response object will include statistics and additional data from the request.
 ### Parameters
 - `model`: (required) the [model name](#model-names)
 - `prompt`: the prompt to generate a response for
 Advanced parameters (optional):
 - `options`: additional model parameters listed in the documentation for the [Modelfile](./modelfile.md#valid-parameters-and-values) such as `temperature`
 - `system`: system prompt to (overrides what is defined in the `Modelfile`)
 - `template`: the full prompt or prompt template (overrides what is defined in the `Modelfile`)
 - `context`: the context parameter returned from a previous request to `/generate`, this can be used to keep a short conversational memory
 - `stream`: if `false` the response will be be returned as a single response object, rather than a stream of objects
 ### Request
 ```shell
 curl -X POST http://localhost:11434/api/generate -d '{
  "model": "llama2:7b",
  "prompt": "Why is the sky blue?"
 }'
 ```
 ### Response
 A stream of JSON objects:
 ```json
 {
  "model": "llama2:7b",
  "created_at": "2023-08-04T08:52:19.385406455-07:00",
  "response": "The",
  "done": false
 }
 ```
 The final response in the stream also includes additional data about the generation:
 - `total_duration`: time spent generating the response
 - `load_duration`: time spent in nanoseconds loading the model
 - `sample_count`: number of samples generated
 - `sample_duration`: time spent generating samples
 - `prompt_eval_count`: number of tokens in the prompt
 - `prompt_eval_duration`: time spent in nanoseconds evaluating the prompt
 - `eval_count`: number of tokens the response
 - `eval_duration`: time in nanoseconds spent generating the response
 - `context`: an encoding of the conversation used in this response, this can be sent in the next request to keep a conversational memory
 - `response`: empty if the response was streamed, if not streamed, this will contain the full response
 To calculate how fast the response is generated in tokens per second (token/s), divide `eval_count` / `eval_duration`.
 ```json
 {
  "model": "llama2:7b",
  "created_at": "2023-08-04T19:22:45.499127Z",
  "response": "",
  "context": [1, 2, 3],
  "done": true,
  "total_duration": 5589157167,
  "load_duration": 3013701500,
  "sample_count": 114,
  "sample_duration": 81442000,
  "prompt_eval_count": 46,
  "prompt_eval_duration": 1160282000,
  "eval_count": 113,
  "eval_duration": 1325948000
 }
 ```
 ## Create a Model
 ```shell
 POST /api/create
 ```
 Create a model from a [`Modelfile`](./modelfile.md)
 ### Parameters
 - `name`: name of the model to create
 - `path`: path to the Modelfile
 - `stream`: (optional) if `false` the response will be be returned as a single response object, rather than a stream of objects
 ### Request
 ```shell
 curl -X POST http://localhost:11434/api/create -d '{
  "name": "mario",
  "path": "~/Modelfile"
 }'
 ```
 ### Response
 A stream of JSON objects. When finished, `status` is `success`.
 ```json
 {
  "status": "parsing modelfile"
 }
 ```
 ## List Local Models
 ```shell
 GET /api/tags
 ```
 List models that are available locally.
 ### Request
 ```shell
 curl http://localhost:11434/api/tags
 ```
 ### Response
 ```json
 {
  "models": [
    {
      "name": "llama2:7b",
      "modified_at": "2023-08-02T17:02:23.713454393-07:00",
      "size": 3791730596
    },
    {
      "name": "llama2:13b",
      "modified_at": "2023-08-08T12:08:38.093596297-07:00",
      "size": 7323310500
    }
  ]
 }
 ```
 ## Show Model Information
 ```shell
 POST /api/show
 ```
 Show details about a model including modelfile, template, parameters, license, and system prompt.
 ### Parameters
 - `name`: name of the model to show
 ### Request
 ```shell
 curl http://localhost:11434/api/show -d '{
  "name": "llama2:7b"
 }'
 ```
 ### Response
 ```json
 {
  "license": "<contents of license block>",
  "modelfile": "# Modelfile generated by \"ollama show\"\n# To build a new Modelfile based on this one, replace the FROM line with:\n# FROM llama2:latest\n\nFROM /Users/username/.ollama/models/blobs/sha256:8daa9615cce30c259a9555b1cc250d461d1bc69980a274b44d7eda0be78076d8\nTEMPLATE \"\"\"[INST] {{ if and .First .System }}<<SYS>>{{ .System }}<</SYS>>\n\n{{ end }}{{ .Prompt }} [/INST] \"\"\"\nSYSTEM \"\"\"\"\"\"\nPARAMETER stop [INST]\nPARAMETER stop [/INST]\nPARAMETER stop <<SYS>>\nPARAMETER stop <</SYS>>\n",
  "parameters": "stop                           [INST]\nstop                           [/INST]\nstop                           <<SYS>>\nstop                           <</SYS>>",
  "template": "[INST] {{ if and .First .System }}<<SYS>>{{ .System }}<</SYS>>\n\n{{ end }}{{ .Prompt }} [/INST] "
 }
 ```
 ## Copy a Model
 ```shell
 POST /api/copy
 ```
 Copy a model. Creates a model with another name from an existing model.
 ### Request
 ```shell
 curl http://localhost:11434/api/copy -d '{
  "source": "llama2:7b",
  "destination": "llama2-backup"
 }'
 ```
 ## Delete a Model
 ```shell
 DELETE /api/delete
 ```
 Delete a model and its data.
 ### Parameters
 - `model`: model name to delete
 ### Request
 ```shell
 curl -X DELETE http://localhost:11434/api/delete -d '{
  "name": "llama2:13b"
 }'
 ```
 ## Pull a Model
 ```shell
 POST /api/pull
 ```
 Download a model from the ollama library. Cancelled pulls are resumed from where they left off, and multiple calls will share the same download progress.
 ### Parameters
 - `name`: name of the model to pull
 - `insecure`: (optional) allow insecure connections to the library. Only use this if you are pulling from your own library during development.
 - `stream`: (optional) if `false` the response will be be returned as a single response object, rather than a stream of objects
 ### Request
 ```shell
 curl -X POST http://localhost:11434/api/pull -d '{
  "name": "llama2:7b"
 }'
 ```
 ### Response
 ```json
 {
  "status": "downloading digestname",
  "digest": "digestname",
  "total": 2142590208
 }
 ```
 ## Push a Model
 ```shell
 POST /api/push
 ```
 Upload a model to a model library. Requires registering for ollama.ai and adding a public key first.
 ### Parameters
 - `name`: name of the model to push in the form of `<namespace>/<model>:<tag>`
 - `insecure`: (optional) allow insecure connections to the library. Only use this if you are pushing to your library during development.
 - `stream`: (optional) if `false` the response will be be returned as a single response object, rather than a stream of objects
 ### Request
 ```shell
 curl -X POST http://localhost:11434/api/push -d '{
  "name": "mattw/pygmalion:latest"
 }'
 ```
 ### Response
 Streaming response that starts with:
 ```json
 { "status": "retrieving manifest" }
 ```
 and then:
 ```json
 {
  "status": "starting upload",
  "digest": "sha256:bc07c81de745696fdf5afca05e065818a8149fb0c77266fb584d9b2cba3711ab",
  "total": 1928429856
 }
 ```
 Then there is a series of uploading responses:
 ```json
 {
  "status": "starting upload",
  "digest": "sha256:bc07c81de745696fdf5afca05e065818a8149fb0c77266fb584d9b2cba3711ab",
  "total": 1928429856
 }
 ```
 Finally, when the upload is complete:
 ```json
 {"status":"pushing manifest"}
 {"status":"success"}
 ```
 ## Generate Embeddings
 ```shell
 POST /api/embeddings
 ```
 Generate embeddings from a model
 ### Parameters
 - `model`: name of model to generate embeddings from
 - `prompt`: text to generate embeddings for
 Advanced parameters:
 - `options`: additional model parameters listed in the documentation for the [Modelfile](./modelfile.md#valid-parameters-and-values) such as `temperature`
 ### Request
 ```shell
 curl -X POST http://localhost:11434/api/embeddings -d '{
  "model": "llama2:7b",
  "prompt": "Here is an article about llamas..."
 }'
 ```
 ### Response
 ```json
 {
  "embeddings": [
    0.5670403838157654, 0.009260174818336964, 0.23178744316101074, -0.2916173040866852, -0.8924556970596313,
    0.8785552978515625, -0.34576427936553955, 0.5742510557174683, -0.04222835972905159, -0.137906014919281
  ]
 }
 ```
--- a/docs/development.md
+++ b/docs/development.md
@@ -1,46 +1,39 @@
 # Development
 - Install cmake or (optionally, required tools for GPUs)
 - run `go generate ./...`
 - run `go build .`
 Install required tools:
-```
+- cmake version 3.24 or higher
-brew install go
+- go version 1.20 or higher
 - gcc version 11.4.0 or higher
 ```bash
 brew install go cmake gcc
 ```
-Enable CGO:
+Get the required libraries:
-```
+```bash
-export CGO_ENABLED=1
+go generate ./...
 ```
 Then build ollama:
-```
+```bash
 go build .
 ```
 Now you can run `ollama`:
-```
+```bash
 ./ollama
 ```
-## Releasing
+## Building on Linux with GPU support
 To release a new version of Ollama you'll need to set some environment variables:
 * `GITHUB_TOKEN`: your GitHub token
 * `APPLE_IDENTITY`: the Apple signing identity (macOS only)
 * `APPLE_ID`: your Apple ID
 * `APPLE_PASSWORD`: your Apple ID app-specific password
 * `APPLE_TEAM_ID`: the Apple team ID for the signing identity
 * `TELEMETRY_WRITE_KEY`: segment write key for telemetry
 Then run the publish script with the target version:
 ```
 VERSION=0.0.2 ./scripts/publish.sh
 ```
 - Install cmake and nvidia-cuda-toolkit
 - run `go generate ./...`
 - run `go build .`
--- a/docs/faq.md
+++ b/docs/faq.md
@@ -0,0 +1,18 @@
 # FAQ
 ## How can I expose the Ollama server?
 ```bash
 OLLAMA_HOST=0.0.0.0:11435 ollama serve
 ```
 By default, Ollama allows cross origin requests from `127.0.0.1` and `0.0.0.0`. To support more origins, you can use the `OLLAMA_ORIGINS` environment variable:
 ```bash
 OLLAMA_ORIGINS=http://192.168.1.1:*,https://example.com ollama serve
 ```
 ## Where are models stored?
 * macOS: Raw model data is stored under `~/.ollama/models`.
 * Linux: Raw model data is stored under `/usr/share/ollama/.ollama/models`
--- a/docs/linux.md
+++ b/docs/linux.md
@@ -0,0 +1,83 @@
 # Installing Ollama on Linux
 > Note: A one line installer for Ollama is available by running:
 >
 > ```bash
 > curl https://ollama.ai/install.sh | sh
 > ```
 ## Download the `ollama` binary
 Ollama is distributed as a self-contained binary. Download it to a directory in your PATH:
 ```bash
 sudo curl -L https://ollama.ai/download/ollama-linux-amd64 -o /usr/bin/ollama
 sudo chmod +x /usr/bin/ollama
 ```
 ## Start Ollama
 Start Ollama by running `ollama serve`:
 ```bash
 ollama serve
 ```
 Once Ollama is running, run a model in another terminal session:
 ```bash
 ollama run llama2
 ```
 ## Install CUDA drivers (optional – for Nvidia GPUs)
 [Download and install](https://developer.nvidia.com/cuda-downloads) CUDA.
 Verify that the drivers are installed by running the following command, which should print details about your GPU:
 ```bash
 nvidia-smi
 ```
 ## Adding Ollama as a startup service (optional)
 Create a user for Ollama:
 ```bash
 sudo useradd -r -s /bin/false -m -d /usr/share/ollama ollama
 ```
 Create a service file in `/etc/systemd/system/ollama.service`:
 ```ini
 [Unit]
 Description=Ollama Service
 After=network-online.target
 [Service]
 ExecStart=/usr/bin/ollama serve
 User=ollama
 Group=ollama
 Restart=always
 RestartSec=3
 Environment="HOME=/usr/share/ollama"
 [Install]
 WantedBy=default.target
 ```
 Then start the service:
 ```bash
 sudo systemctl daemon-reload
 sudo systemctl enable ollama
 ```
 ### Viewing logs
 To view logs of Ollama running as a startup service, run:
 ```bash
 journalctl -u ollama
 ```
--- a/docs/modelfile.md
+++ b/docs/modelfile.md
@@ -1,105 +1,191 @@
 # Ollama Model File
-> Note: this model file syntax is in development
+> Note: this `Modelfile` syntax is in development
 A model file is the blueprint to create and share models with Ollama.
 ## Table of Contents
 - [Format](#format)
 - [Examples](#examples)
 - [Instructions](#instructions)
  - [FROM (Required)](#from-required)
    - [Build from llama2](#build-from-llama2)
    - [Build from a bin file](#build-from-a-bin-file)
  - [EMBED](#embed)
  - [PARAMETER](#parameter)
    - [Valid Parameters and Values](#valid-parameters-and-values)
  - [TEMPLATE](#template)
    - [Template Variables](#template-variables)
  - [SYSTEM](#system)
  - [ADAPTER](#adapter)
  - [LICENSE](#license)
 - [Notes](#notes)
 ## Format
-The format of the Modelfile:
+The format of the `Modelfile`:
 ```modelfile
 # comment
 INSTRUCTION arguments
 ```
-| Instruction       | Description                                           |
+| Instruction                         | Description                                                   |
-| ----------------- | ----------------------------------------------------- |
+| ----------------------------------- | ------------------------------------------------------------- |
-| `FROM` (required) | Defines the base model to use                         |
+| [`FROM`](#from-required) (required) | Defines the base model to use.                                |
-| `PARAMETER`       | Sets the parameters for how Ollama will run the model |
+| [`PARAMETER`](#parameter)           | Sets the parameters for how Ollama will run the model.        |
-| `SYSTEM`          | Specifies the system prompt that will set the context |
+| [`TEMPLATE`](#template)             | The full prompt template to be sent to the model.             |
-| `TEMPLATE`        | The full prompt template to be sent to the model      |
+| [`SYSTEM`](#system)                 | Specifies the system prompt that will be set in the template. |
-| `LICENSE`         | Specifies the legal license                           |
+| [`ADAPTER`](#adapter)               | Defines the (Q)LoRA adapters to apply to the model.           |
 | [`LICENSE`](#license)               | Specifies the legal license.                                  |
 ## Examples
-An example of a model file creating a mario blueprint:
+An example of a `Modelfile` creating a mario blueprint:
-```
+```modelfile
 FROM llama2
 # sets the temperature to 1 [higher is more creative, lower is more coherent]
 # sets the context size to 4096
 PARAMETER temperature 1
 # sets the context window size to 4096, this controls how many tokens the LLM can use as context to generate the next token
 PARAMETER num_ctx 4096
-# Overriding the system prompt
+# sets a custom system prompt to specify the behavior of the chat assistant
 SYSTEM You are Mario from super mario bros, acting as an assistant.
 ```
 To use this:
-1. Save it as a file (eg. `Modelfile`)
+1. Save it as a file (e.g. `Modelfile`)
-2. `ollama create NAME -f <location of the file eg. ./Modelfile>'`
+2. `ollama create choose-a-model-name -f <location of the file e.g. ./Modelfile>'`
-3. `ollama run NAME`
+3. `ollama run choose-a-model-name`
 4. Start using the model!
-## FROM (Required)
+More examples are available in the [examples directory](../examples).
-The FROM instruction defines the base model to use when creating a model.
+## Instructions
-```
+### FROM (Required)
 The `FROM` instruction defines the base model to use when creating a model.
 ```modelfile
 FROM <model name>:<tag>
 ```
-### Build from llama2
+#### Build from llama2
-```
+```modelfile
 FROM llama2
 ```
 A list of available base models:
 <https://github.com/jmorganca/ollama#model-library>
-### Build from a bin file
+#### Build from a `bin` file
-```
+```modelfile
 FROM ./ollama-model.bin
 ```
-## PARAMETER (Optional)
+This bin file location should be specified as an absolute path or relative to the `Modelfile` location.
 ### EMBED
 The `EMBED` instruction is used to add embeddings of files to a model. This is useful for adding custom data that the model can reference when generating an answer. Note that currently only text files are supported, formatted with each line as one embedding.
 ```modelfile
 FROM <model name>:<tag>
 EMBED <file path>.txt
 EMBED <different file path>.txt
 EMBED <path to directory>/*.txt
 ```
 ### PARAMETER
 The `PARAMETER` instruction defines a parameter that can be set when the model is run.
-```
+```modelfile
 PARAMETER <parameter> <parametervalue>
 ```
 ### Valid Parameters and Values
-| Parameter      | Description                                                                                                                                                                                                                                             | Value Type | Example Usage      |
+| Parameter      | Description                                                                                                                                                                                                                                             | Value Type | Example Usage        |
-| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | ------------------ |
+| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | -------------------- |
-| num_ctx        | Sets the size of the prompt context size length model. (Default: 2048)                                                                                                                                                                                  | int        | num_ctx 4096       |
+| mirostat       | Enable Mirostat sampling for controlling perplexity. (default: 0, 0 = disabled, 1 = Mirostat, 2 = Mirostat 2.0)                                                                                                                                         | int        | mirostat 0           |
-| temperature    | The temperature of the model. Increasing the temperature will make the model answer more creatively. (Default: 0.8)                                                                                                                                     | float      | temperature 0.7    |
+| mirostat_eta   | Influences how quickly the algorithm responds to feedback from the generated text. A lower learning rate will result in slower adjustments, while a higher learning rate will make the algorithm more responsive. (Default: 0.1)                        | float      | mirostat_eta 0.1     |
-| top_k          | Reduces the probability of generating nonsense. A higher value (e.g. 100) will give more diverse answers, while a lower value (e.g. 10) will be more conservative. (Default: 40)                                                                        | int        | top_k 40           |
+| mirostat_tau   | Controls the balance between coherence and diversity of the output. A lower value will result in more focused and coherent text. (Default: 5.0)                                                                                                         | float      | mirostat_tau 5.0     |
-| top_p          | Works together with top-k. A higher value (e.g., 0.95) will lead to more diverse text, while a lower value (e.g., 0.5) will generate more focused and conservative text. (Default: 0.9)                                                                 | float      | top_p 0.9          |
+| num_ctx        | Sets the size of the context window used to generate the next token. (Default: 2048)                                                                                                                                                                    | int        | num_ctx 4096         |
-| num_gpu        | The number of GPUs to use. On macOS it defaults to 1 to enable metal support, 0 to disable.                                                                                                                                                             | int        | num_gpu 1          |
+| num_gqa        | The number of GQA groups in the transformer layer. Required for some models, for example it is 8 for llama2:70b                                                                                                                                         | int        | num_gqa 1            |
-| repeat_last_n  | Sets how far back for the model to look back to prevent repetition. (Default: 64, 0 = disabled, -1 = ctx-size)                                                                                                                                          | int        | repeat_last_n 64   |
+| num_gpu        | The number of layers to send to the GPU(s). On macOS it defaults to 1 to enable metal support, 0 to disable.                                                                                                                                            | int        | num_gpu 50           |
-| repeat_penalty | Sets how strongly to penalize repetitions. A higher value (e.g., 1.5) will penalize repetitions more strongly, while a lower value (e.g., 0.9) will be more lenient. (Default: 1.1)                                                                     | float      | repeat_penalty 1.1 |
+| num_thread     | Sets the number of threads to use during computation. By default, Ollama will detect this for optimal performance. It is recommended to set this value to the number of physical CPU cores your system has (as opposed to the logical number of cores). | int        | num_thread 8         |
-| tfs_z          | Tail free sampling is used to reduce the impact of less probable tokens from the output. A higher value (e.g., 2.0) will reduce the impact more, while a value of 1.0 disables this setting. (default: 1)                                               | float      | tfs_z 1            |
+| repeat_last_n  | Sets how far back for the model to look back to prevent repetition. (Default: 64, 0 = disabled, -1 = num_ctx)                                                                                                                                           | int        | repeat_last_n 64     |
-| mirostat       | Enable Mirostat sampling for controlling perplexity. (default: 0, 0 = disabled, 1 = Mirostat, 2 = Mirostat 2.0)                                                                                                                                         | int        | mirostat 0         |
+| repeat_penalty | Sets how strongly to penalize repetitions. A higher value (e.g., 1.5) will penalize repetitions more strongly, while a lower value (e.g., 0.9) will be more lenient. (Default: 1.1)                                                                     | float      | repeat_penalty 1.1   |
-| mirostat_tau   | Controls the balance between coherence and diversity of the output. A lower value will result in more focused and coherent text. (Default: 5.0)                                                                                                         | float      | mirostat_tau 5.0   |
+| temperature    | The temperature of the model. Increasing the temperature will make the model answer more creatively. (Default: 0.8)                                                                                                                                     | float      | temperature 0.7      |
-| mirostat_eta   | Influences how quickly the algorithm responds to feedback from the generated text. A lower learning rate will result in slower adjustments, while a higher learning rate will make the algorithm more responsive. (Default: 0.1)                        | float      | mirostat_eta 0.1   |
+| seed | Sets the random number seed to use for generation. Setting this to a specific number will make the model generate the same text for the same prompt. | int | seed 42 |
-| num_thread     | Sets the number of threads to use during computation. By default, Ollama will detect this for optimal performance. It is recommended to set this value to the number of physical CPU cores your system has (as opposed to the logical number of cores). | int        | num_thread 8       |
+| stop           | Sets the stop sequences to use.                                                                                                                                                                                                                         | string     | stop "AI assistant:" |
 | tfs_z          | Tail free sampling is used to reduce the impact of less probable tokens from the output. A higher value (e.g., 2.0) will reduce the impact more, while a value of 1.0 disables this setting. (default: 1)                                               | float      | tfs_z 1              |
 | num_predict    | Maximum number of tokens to predict when generating text. (Default: 128, -1 = infinite generation, -2 = fill context)                                                                                                                                   | int        | num_predict 42       |
 | top_k          | Reduces the probability of generating nonsense. A higher value (e.g. 100) will give more diverse answers, while a lower value (e.g. 10) will be more conservative. (Default: 40)                                                                        | int        | top_k 40             |
 | top_p          | Works together with top-k. A higher value (e.g., 0.95) will lead to more diverse text, while a lower value (e.g., 0.5) will generate more focused and conservative text. (Default: 0.9)                                                                 | float      | top_p 0.9            |
-## Prompt
+### TEMPLATE
-When building on top of the base models supplied by Ollama, it comes with the prompt template predefined. To override the supplied system prompt, simply add `SYSTEM insert system prompt` to change the system prompt.
+`TEMPLATE` of the full prompt template to be passed into the model. It may include (optionally) a system prompt and a user's prompt. This is used to create a full custom prompt, and syntax may be model specific. You can usually find the template for a given model in the readme for that model.
-### Prompt Template
+#### Template Variables
-`TEMPLATE` the full prompt template to be passed into the model. It may include (optionally) a system prompt, user prompt, and assistant prompt. This is used to create a full custom prompt, and syntax may be model specific.
+| Variable        | Description                                                                                                  |
 | --------------- | ------------------------------------------------------------------------------------------------------------ |
 | `{{ .System }}` | The system prompt used to specify custom behavior, this must also be set in the Modelfile as an instruction. |
 | `{{ .Prompt }}` | The incoming prompt, this is not specified in the model file and will be set based on input.                 |
 | `{{ .First }}`  | A boolean value used to render specific template information for the first generation of a session.          |
 ```modelfile
 TEMPLATE """
 {{- if .First }}
 ### System:
 {{ .System }}
 {{- end }}
 ### User:
 {{ .Prompt }}
 ### Response:
 """
 SYSTEM """<system message>"""
 ```
 ### SYSTEM
 The `SYSTEM` instruction specifies the system prompt to be used in the template, if applicable.
 ```modelfile
 SYSTEM """<system message>"""
 ```
 ### ADAPTER
 The `ADAPTER` instruction specifies the LoRA adapter to apply to the base model. The value of this instruction should be an absolute path or a path relative to the Modelfile and the file must be in a GGML file format. The adapter should be tuned from the base model otherwise the behaviour is undefined.
 ```modelfile
 ADAPTER ./ollama-lora.bin
 ```
 ### LICENSE
 The `LICENSE` instruction allows you to specify the legal license under which the model used with this Modelfile is shared or distributed.
 ```modelfile
 LICENSE """
 <license text>
 """
 ```
 ## Notes
- the **modelfile is not case sensitive**. In the examples, we use uppercase for instructions to make it easier to distinguish it from arguments.
+- the **`Modelfile` is not case sensitive**. In the examples, we use uppercase for instructions to make it easier to distinguish it from arguments.
 - Instructions can be in any order. In the examples, we start with FROM instruction to keep it easily readable.
--- a/docs/quantize.md
+++ b/docs/quantize.md
@@ -0,0 +1,111 @@
 # How to Quantize a Model
 Sometimes the model you want to work with is not available at [https://ollama.ai/library](https://ollama.ai/library).
 ## Figure out if we can run the model?
 Not all models will work with Ollama. There are a number of factors that go into whether we are able to work with the next cool model. First it has to work with llama.cpp. Then we have to have implemented the features of llama.cpp that it requires. And then, sometimes, even with both of those, the model might not work...
 1. What is the model you want to convert and upload?
 2. Visit the model's page on HuggingFace.
 3. Switch to the **Files and versions** tab.
 4. Click on the **config.json** file. If there is no config.json file, it may not work.
 5. Take note of the **architecture** list in the json file.
 6. Does any entry in the list match one of the following architectures?
    1. LlamaForCausalLM
    2. MistralForCausalLM
    3. RWForCausalLM
    4. FalconForCausalLM
    5. GPTNeoXForCausalLM
    6. GPTBigCodeForCausalLM
 7. If the answer is yes, then there is a good chance the model will run after being converted and quantized.
 8. An alternative to this process is to visit [https://caniquant.tvl.st](https://caniquant.tvl.st) and enter the org/modelname in the box and submit.
 At this point there are two processes you can use. You can either use a Docker container to convert and quantize, OR you can manually run the scripts. The Docker container is the easiest way to do it, but it requires you to have Docker installed on your machine. If you don't have Docker installed, you can follow the manual process.
 ## Convert and Quantize with Docker
 Run `docker run --rm -v /path/to/model/repo:/repo ollama/quantize -q quantlevel /repo`. For instance, if you have downloaded the latest Mistral 7B model, then clone it to your machine. Then change into that directory and you can run:
 ```shell
 docker run --rm -v .:/repo ollama/quantize -q q4_0 /repo
 ```
 You can find the different quantization levels below under **Quantize the Model**.
 This will output two files into the directory. First is a f16.bin file that is the model converted to GGUF. The second file is a q4_0.bin file which is the model quantized to a 4 bit quantization. You should rename it to something more descriptive.
 You can find the repository for the Docker container here: [https://github.com/mxyng/quantize](https://github.com/mxyng/quantize)
 For instance, if you wanted to convert the Mistral 7B model to a Q4 quantized model, then you could go through the following steps:
 1. First verify the model will potentially work.
 2. Now clone Mistral 7B to your machine. You can find the command to run when you click the three vertical dots button on the model page, then click **Clone Repository**.
   1. For this repo, the command is:
      ```shell
      git lfs install
      git clone https://huggingface.co/mistralai/Mistral-7B-v0.1
      ```
   2. Navigate into the new directory and run `docker run --rm -v .:/repo ollama/quantize -q q4_0 /repo`
   3. Now you can create a modelfile using the q4_0.bin file that was created.
 ## Convert and Quantize Manually
 ### Clone llama.cpp to your machine
 If we know the model has a chance of working, then we need to convert and quantize. This is a matter of running two separate scripts in the llama.cpp project.
 1. Decide where you want the llama.cpp repository on your machine.
 2. Navigate to that location and then run:
 [`git clone https://github.com/ggerganov/llama.cpp.git`](https://github.com/ggerganov/llama.cpp.git)
    1. If you don't have git installed, download this zip file and unzip it to that location: https://github.com/ggerganov/llama.cpp/archive/refs/heads/master.zip
 3. Install the Python dependencies: `pip install torch transformers sentencepiece`
 4. Run 'make' to build the project and the quantize executable.
 ### Convert the model to GGUF
 1. Decide on the right convert script to run. What was the model architecture you found in the first section.
    1. LlamaForCausalLM or MistralForCausalLM:
    run `python3 convert.py <modelfilename>`
    No need to specify fp16 or fp32.
    2. FalconForCausalLM or RWForCausalLM:
    run `python3 convert-falcon-hf-to-gguf.py <modelfilename> <fpsize>`  
    fpsize depends on the weight size. 1 for fp16, 0 for fp32
    3. GPTNeoXForCausalLM:
    run `python3 convert-gptneox-hf-to-gguf.py <modelfilename> <fpsize>`
    fpsize depends on the weight size. 1 for fp16, 0 for fp32
    4. GPTBigCodeForCausalLM:
    run `python3 convert-starcoder-hf-to-gguf.py <modelfilename> <fpsize>`
    fpsize depends on the weight size. 1 for fp16, 0 for fp32
 ### Quantize the model
 If the model converted successfully, there is a good chance it will also quantize successfully. Now you need to decide on the quantization to use. We will always try to create all the quantizations and upload them to the library. You should decide which level is more important to you and quantize accordingly.
 The quantization options are as follows. Note that some architectures such as Falcon do not support K quants.
 - Q4_0
 - Q4_1
 - Q5_0
 - Q5_1
 - Q2_K
 - Q3_K
 - Q3_K_S
 - Q3_K_M
 - Q3_K_L
 - Q4_K
 - Q4_K_S
 - Q4_K_M
 - Q5_K
 - Q5_K_S
 - Q5_K_M
 - Q6_K
 - Q8_0
 Run the following command `quantize <converted model from above> <output file> <quantization type>`
 ## Now Create the Model
 Now you can create the Ollama model. Refer to the [modelfile](./modelfile.md) doc for more information on doing that.
--- a/docs/tutorials.md
+++ b/docs/tutorials.md
@@ -0,0 +1,8 @@
 # Tutorials
 Here is a list of ways you can use Ollama with other tools to build interesting applications.
 - [Using LangChain with Ollama in JavaScript](./tutorials/langchainjs.md)
 - [Using LangChain with Ollama in Python](./tutorials/langchainpy.md)
 Also be sure to check out the [examples](../examples) directory for more ways to use Ollama.
--- a/docs/tutorials/langchainjs.md
+++ b/docs/tutorials/langchainjs.md
@@ -0,0 +1,73 @@
 # Using LangChain with Ollama using JavaScript
 In this tutorial, we are going to use JavaScript with LangChain and Ollama to learn about something just a touch more recent. In August 2023, there was a series of wildfires on Maui. There is no way an LLM trained before that time can know about this, since their training data would not include anything as recent as that. So we can find the [Wikipedia article about the fires](https://en.wikipedia.org/wiki/2023_Hawaii_wildfires) and ask questions about the contents.
 To get started, let's just use **LangChain** to ask a simple question to a model. To do this with JavaScript, we need to install **LangChain**:
 ```bash
 npm install langchain
 ```
 Now we can start building out our JavaScript:
 ```javascript
 import { Ollama } from "langchain/llms/ollama";
 const ollama = new Ollama({
  baseUrl: "http://localhost:11434",
  model: "llama2",
 });
 const answer = await ollama.call(`why is the sky blue?`);
 console.log(answer);
 ```
 That will get us the same thing as if we ran `ollama run llama2 "why is the sky blue"` in the terminal. But we want to load a document from the web to ask a question against. **Cheerio** is a great library for ingesting a webpage, and **LangChain** uses it in their **CheerioWebBaseLoader**. So let's build that part of the app.
 ```javascript
 import { CheerioWebBaseLoader } from "langchain/document_loaders/web/cheerio";
 const loader = new CheerioWebBaseLoader("https://en.wikipedia.org/wiki/2023_Hawaii_wildfires");
 const data = loader.load();
 ```
 That will load the document. Although this page is smaller than the Odyssey, it is certainly bigger than the context size for most LLMs. So we are going to need to split into smaller pieces, and then select just the pieces relevant to our question. This is a great use for a vector datastore. In this example, we will use the **MemoryVectorStore** that is part of **LangChain**. But there is one more thing we need to get the content into the datastore. We have to run an embeddings process that converts the tokens in the text into a series of vectors. And for that, we are going to use **Tensorflow**. There is a lot of stuff going on in this one. First, install the **Tensorflow** components that we need.
 ```javascript
 npm install @tensorflow/tfjs-core@3.6.0 @tensorflow/tfjs-converter@3.6.0 @tensorflow-models/universal-sentence-encoder@1.3.3 @tensorflow/tfjs-node@4.10.0
 ```
 If you just install those components without the version numbers, it will install the latest versions, but there are conflicts within **Tensorflow**, so you need to install the compatible versions.
 ```javascript
 import { RecursiveCharacterTextSplitter } from "langchain/text_splitter"
 import { MemoryVectorStore } from "langchain/vectorstores/memory";
 import "@tensorflow/tfjs-node";
 import { TensorFlowEmbeddings } from "langchain/embeddings/tensorflow";
 // Split the text into 500 character chunks. And overlap each chunk by 20 characters
 const textSplitter = new RecursiveCharacterTextSplitter({
 chunkSize: 500,
 chunkOverlap: 20
 });
 const splitDocs = await textSplitter.splitDocuments(data);
 // Then use the TensorFlow Embedding to store these chunks in the datastore
 const vectorStore = await MemoryVectorStore.fromDocuments(splitDocs, new TensorFlowEmbeddings());
 ```
 To connect the datastore to a question asked to a LLM, we need to use the concept at the heart of **LangChain**: the chain. Chains are a way to connect a number of activities together to accomplish a particular tasks. There are a number of chain types available, but for this tutorial we are using the **RetrievalQAChain**.
 ```javascript
 import { RetrievalQAChain } from "langchain/chains";
 const retriever = vectorStore.asRetriever();
 const chain = RetrievalQAChain.fromLLM(ollama, retriever);
 const result = await chain.call({query: "When was Hawaii's request for a major disaster declaration approved?"});
 console.log(result.text)
 ```
 So we created a retriever, which is a way to return the chunks that match a query from a datastore. And then connect the retriever and the model via a chain. Finally, we send a query to the chain, which results in an answer using our document as a source. The answer it returned was correct, August 10, 2023.
 And that is a simple introduction to what you can do with **LangChain** and **Ollama.**
--- a/docs/tutorials/langchainpy.md
+++ b/docs/tutorials/langchainpy.md
@@ -0,0 +1,81 @@
 # Using LangChain with Ollama in Python
 Let's imagine we are studying the classics, such as **the Odyssey** by **Homer**. We might have a question about Neleus and his family. If you ask llama2 for that info, you may get something like:
 > I apologize, but I'm a large language model, I cannot provide information on individuals or families that do not exist in reality. Neleus is not a real person or character, and therefore does not have a family or any other personal details. My apologies for any confusion. Is there anything else I can help you with?
 This sounds like a typical censored response, but even llama2-uncensored gives a mediocre answer:
 > Neleus was a legendary king of Pylos and the father of Nestor, one of the Argonauts. His mother was Clymene, a sea nymph, while his father was Neptune, the god of the sea.
 So let's figure out how we can use **LangChain** with Ollama to ask our question to the actual document, the Odyssey by Homer, using Python.
 Let's start by asking a simple question that we can get an answer to from the **Llama2** model using **Ollama**. First, we need to install the **LangChain** package:
 `pip install langchain`
 Then we can create a model and ask the question:
 ```python
 from langchain.llms import Ollama
 ollama = Ollama(base_url='http://localhost:11434',
 model="llama2")
 print(ollama("why is the sky blue"))
 ```
 Notice that we are defining the model and the base URL for Ollama.
 Now let's load a document to ask questions against. I'll load up the Odyssey by Homer, which you can find at Project Gutenberg. We will need **WebBaseLoader** which is part of **LangChain** and loads text from any webpage. On my machine, I also needed to install **bs4** to get that to work, so run `pip install bs4`.
 ```python
 from langchain.document_loaders import WebBaseLoader
 loader = WebBaseLoader("https://www.gutenberg.org/files/1727/1727-h/1727-h.htm")
 data = loader.load()
 ```
 This file is pretty big. Just the preface is 3000 tokens. Which means the full document won't fit into the context for the model. So we need to split it up into smaller pieces.
 ```python
 from langchain.text_splitter import RecursiveCharacterTextSplitter
 text_splitter=RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=0)
 all_splits = text_splitter.split_documents(data)
 ```
 It's split up, but we have to find the relevant splits and then submit those to the model. We can do this by creating embeddings and storing them in a vector database. For now, we don't have embeddings built in to Ollama, though we will be adding that soon, so for now, we can use the GPT4All library for that. We will use ChromaDB in this example for a vector database. `pip install GPT4All chromadb`
 ```python
 from langchain.embeddings import GPT4AllEmbeddings
 from langchain.vectorstores import Chroma
 vectorstore = Chroma.from_documents(documents=all_splits, embedding=GPT4AllEmbeddings())
 ```
 Now let's ask a question from the document. **Who was Neleus, and who is in his family?** Neleus is a character in the Odyssey, and the answer can be found in our text.
 ```python
 question="Who is Neleus and who is in Neleus' family?"
 docs = vectorstore.similarity_search(question)
 len(docs)
 ```
 This will output the number of matches for chunks of data similar to the search.
 The next thing is to send the question and the relevant parts of the docs to the model to see if we can get a good answer. But we are stitching two parts of the process together, and that is called a chain. This means we need to define a chain:
 ```python
 from langchain.chains import RetrievalQA
 qachain=RetrievalQA.from_chain_type(ollama, retriever=vectorstore.as_retriever())
 qachain({"query": question})
 ```
 The answer received from this chain was:
 > Neleus is a character in Homer's "Odyssey" and is mentioned in the context of Penelope's suitors. Neleus is the father of Chloris, who is married to Neleus and bears him several children, including Nestor, Chromius, Periclymenus, and Pero. Amphinomus, the son of Nisus, is also mentioned as a suitor of Penelope and is known for his good natural disposition and agreeable conversation.
 It's not a perfect answer, as it implies Neleus married his daughter when actually Chloris "was the youngest daughter to Amphion son of Iasus and king of Minyan Orchomenus, and was Queen in Pylos".
 I updated the chunk_overlap for the text splitter to 20 and tried again and got a much better answer:
 > Neleus is a character in Homer's epic poem "The Odyssey." He is the husband of Chloris, who is the youngest daughter of Amphion son of Iasus and king of Minyan Orchomenus. Neleus has several children with Chloris, including Nestor, Chromius, Periclymenus, and Pero.
 And that is a much better answer.
--- a/examples/.gitignore
+++ b/examples/.gitignore
@@ -0,0 +1,171 @@
 node_modules
 # OSX
 .DS_STORE
 # Models
 models/
 # Local Chroma db
 .chroma/
 db/
 # Byte-compiled / optimized / DLL files
 __pycache__/
 *.py[cod]
 *$py.class
 # C extensions
 *.so
 # Distribution / packaging
 .Python
 build/
 develop-eggs/
 dist/
 downloads/
 eggs/
 .eggs/
 lib/
 lib64/
 parts/
 sdist/
 var/
 wheels/
 share/python-wheels/
 *.egg-info/
 .installed.cfg
 *.egg
 MANIFEST
 # PyInstaller
 #  Usually these files are written by a python script from a template
 #  before PyInstaller builds the exe, so as to inject date/other infos into it.
 *.manifest
 *.spec
 # Installer logs
 pip-log.txt
 pip-delete-this-directory.txt
 # Unit test / coverage reports
 htmlcov/
 .tox/
 .nox/
 .coverage
 .coverage.*
 .cache
 nosetests.xml
 coverage.xml
 *.cover
 *.py,cover
 .hypothesis/
 .pytest_cache/
 cover/
 # Translations
 *.mo
 *.pot
 # Django stuff:
 *.log
 local_settings.py
 db.sqlite3
 db.sqlite3-journal
 # Flask stuff:
 instance/
 .webassets-cache
 # Scrapy stuff:
 .scrapy
 # Sphinx documentation
 docs/_build/
 # PyBuilder
 .pybuilder/
 target/
 # Jupyter Notebook
 .ipynb_checkpoints
 # IPython
 profile_default/
 ipython_config.py
 # pyenv
 #   For a library or package, you might want to ignore these files since the code is
 #   intended to run in multiple environments; otherwise, check them in:
 # .python-version
 # pipenv
 #   According to pypa/pipenv#598, it is recommended to include Pipfile.lock in version control.
 #   However, in case of collaboration, if having platform-specific dependencies or dependencies
 #   having no cross-platform support, pipenv may install dependencies that don't work, or not
 #   install all needed dependencies.
 #Pipfile.lock
 # poetry
 #   Similar to Pipfile.lock, it is generally recommended to include poetry.lock in version control.
 #   This is especially recommended for binary packages to ensure reproducibility, and is more
 #   commonly ignored for libraries.
 #   https://python-poetry.org/docs/basic-usage/#commit-your-poetrylock-file-to-version-control
 #poetry.lock
 # pdm
 #   Similar to Pipfile.lock, it is generally recommended to include pdm.lock in version control.
 #pdm.lock
 #   pdm stores project-wide configurations in .pdm.toml, but it is recommended to not include it
 #   in version control.
 #   https://pdm.fming.dev/#use-with-ide
 .pdm.toml
 # PEP 582; used by e.g. github.com/David-OConnor/pyflow and github.com/pdm-project/pdm
 __pypackages__/
 # Celery stuff
 celerybeat-schedule
 celerybeat.pid
 # SageMath parsed files
 *.sage.py
 # Environments
 .env
 .venv
 env/
 venv/
 ENV/
 env.bak/
 venv.bak/
 # Spyder project settings
 .spyderproject
 .spyproject
 # Rope project settings
 .ropeproject
 # mkdocs documentation
 /site
 # mypy
 .mypy_cache/
 .dmypy.json
 dmypy.json
 # Pyre type checker
 .pyre/
 # pytype static type analyzer
 .pytype/
 # Cython debug symbols
 cython_debug/
 # PyCharm
 #  JetBrains specific template is maintained in a separate JetBrains.gitignore that can
 #  be found at https://github.com/github/gitignore/blob/main/Global/JetBrains.gitignore
 #  and can be added to the global gitignore or merged into this file.  For a more nuclear
 #  option (not recommended) you can uncomment the following to ignore the entire idea folder.
 #.idea/
--- a/examples/README.md
+++ b/examples/README.md
@@ -1,15 +1,3 @@
 # Examples
-This directory contains examples that can be created and run with `ollama`.
+This directory contains different examples of using Ollama.
 To create a model:
 ```
 ollama create example -f <example file>
 ```
 To run a model:
 ```
 ollama run example
 ```
--- a/examples/golang-simplegenerate/README.md
+++ b/examples/golang-simplegenerate/README.md
--- a/examples/golang-simplegenerate/main.go
+++ b/examples/golang-simplegenerate/main.go
@@ -0,0 +1,27 @@
 package main
 import (
 	"bytes"
 	"fmt"
 	"net/http"
 	"os"
 	"io"
 	"log"
 )
 func main() {
 	body := []byte(`{"model":"mistral"}`)
 	resp, err := http.Post("http://localhost:11434/api/generate", "application/json", bytes.NewBuffer(body))
 	if err != nil {
 		fmt.Print(err.Error())
 		os.Exit(1)
 	} 
 	responseData, err := io.ReadAll(resp.Body)
 	if err != nil {
 		log.Fatal(err)
 	}
 	fmt.Println(string(responseData))
 }
--- a/examples/langchain-python-rag-document/README.md
+++ b/examples/langchain-python-rag-document/README.md
@@ -0,0 +1,21 @@
 # LangChain Document QA
 This example provides an interface for asking questions to a PDF document.
 ## Setup
 ```
 pip install -r requirements.txt
 ```
 ## Run
 ```
 python main.py
 ```
 A prompt will appear, where questions may be asked:
 ```
 Query: How many locations does WeWork have?
 ```
--- a/examples/langchain-python-rag-document/main.py
+++ b/examples/langchain-python-rag-document/main.py
@@ -0,0 +1,61 @@
 from langchain.document_loaders import OnlinePDFLoader
 from langchain.vectorstores import Chroma
 from langchain.embeddings import GPT4AllEmbeddings
 from langchain import PromptTemplate
 from langchain.llms import Ollama
 from langchain.callbacks.manager import CallbackManager
 from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
 from langchain.chains import RetrievalQA
 import sys
 import os
 class SuppressStdout:
    def __enter__(self):
        self._original_stdout = sys.stdout
        self._original_stderr = sys.stderr
        sys.stdout = open(os.devnull, 'w')
        sys.stderr = open(os.devnull, 'w')
    def __exit__(self, exc_type, exc_val, exc_tb):
        sys.stdout.close()
        sys.stdout = self._original_stdout
        sys.stderr = self._original_stderr
 # load the pdf and split it into chunks
 loader = OnlinePDFLoader("https://d18rn0p25nwr6d.cloudfront.net/CIK-0001813756/975b3e9b-268e-4798-a9e4-2a9a7c92dc10.pdf")
 data = loader.load()
 from langchain.text_splitter import RecursiveCharacterTextSplitter
 text_splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=0)
 all_splits = text_splitter.split_documents(data)
 with SuppressStdout():
    vectorstore = Chroma.from_documents(documents=all_splits, embedding=GPT4AllEmbeddings())
 while True:
    query = input("\nQuery: ")
    if query == "exit":
        break
    if query.strip() == "":
        continue
    # Prompt
    template = """Use the following pieces of context to answer the question at the end. 
    If you don't know the answer, just say that you don't know, don't try to make up an answer. 
    Use three sentences maximum and keep the answer as concise as possible. 
    {context}
    Question: {question}
    Helpful Answer:"""
    QA_CHAIN_PROMPT = PromptTemplate(
        input_variables=["context", "question"],
        template=template,
    )
    llm = Ollama(model="llama2:13b", callback_manager=CallbackManager([StreamingStdOutCallbackHandler()]))
    qa_chain = RetrievalQA.from_chain_type(
        llm,
        retriever=vectorstore.as_retriever(),
        chain_type_kwargs={"prompt": QA_CHAIN_PROMPT},
    )
    result = qa_chain({"query": query})
--- a/examples/langchain-python-rag-document/requirements.txt
+++ b/examples/langchain-python-rag-document/requirements.txt
@@ -0,0 +1,109 @@
 absl-py==1.4.0
 aiohttp==3.8.5
 aiosignal==1.3.1
 anyio==3.7.1
 astunparse==1.6.3
 async-timeout==4.0.3
 attrs==23.1.0
 backoff==2.2.1
 beautifulsoup4==4.12.2
 bs4==0.0.1
 cachetools==5.3.1
 certifi==2023.7.22
 cffi==1.15.1
 chardet==5.2.0
 charset-normalizer==3.2.0
 Chroma==0.2.0
 chroma-hnswlib==0.7.2
 chromadb==0.4.5
 click==8.1.6
 coloredlogs==15.0.1
 cryptography==41.0.3
 dataclasses-json==0.5.14
 fastapi==0.99.1
 filetype==1.2.0
 flatbuffers==23.5.26
 frozenlist==1.4.0
 gast==0.4.0
 google-auth==2.22.0
 google-auth-oauthlib==1.0.0
 google-pasta==0.2.0
 gpt4all==1.0.8
 grpcio==1.57.0
 h11==0.14.0
 h5py==3.9.0
 httptools==0.6.0
 humanfriendly==10.0
 idna==3.4
 importlib-resources==6.0.1
 joblib==1.3.2
 keras==2.13.1
 langchain==0.0.261
 langsmith==0.0.21
 libclang==16.0.6
 lxml==4.9.3
 Markdown==3.4.4
 MarkupSafe==2.1.3
 marshmallow==3.20.1
 monotonic==1.6
 mpmath==1.3.0
 multidict==6.0.4
 mypy-extensions==1.0.0
 nltk==3.8.1
 numexpr==2.8.5
 numpy==1.24.3
 oauthlib==3.2.2
 onnxruntime==1.15.1
 openapi-schema-pydantic==1.2.4
 opt-einsum==3.3.0
 overrides==7.4.0
 packaging==23.1
 pdf2image==1.16.3
 pdfminer==20191125
 pdfminer.six==20221105
 Pillow==10.0.0
 posthog==3.0.1
 protobuf==4.24.0
 pulsar-client==3.2.0
 pyasn1==0.5.0
 pyasn1-modules==0.3.0
 pycparser==2.21
 pycryptodome==3.18.0
 pydantic==1.10.12
 PyPika==0.48.9
 python-dateutil==2.8.2
 python-dotenv==1.0.0
 python-magic==0.4.27
 PyYAML==6.0.1
 regex==2023.8.8
 requests==2.31.0
 requests-oauthlib==1.3.1
 rsa==4.9
 six==1.16.0
 sniffio==1.3.0
 soupsieve==2.4.1
 SQLAlchemy==2.0.19
 starlette==0.27.0
 sympy==1.12
 tabulate==0.9.0
 tenacity==8.2.2
 tensorboard==2.13.0
 tensorboard-data-server==0.7.1
 tensorflow==2.13.0
 tensorflow-estimator==2.13.0
 tensorflow-hub==0.14.0
 tensorflow-macos==2.13.0
 termcolor==2.3.0
 tokenizers==0.13.3
 tqdm==4.66.1
 typing-inspect==0.9.0
 typing_extensions==4.5.0
 unstructured==0.9.2
 urllib3==1.26.16
 uvicorn==0.23.2
 uvloop==0.17.0
 watchfiles==0.19.0
 websockets==11.0.3
 Werkzeug==2.3.6
 wrapt==1.15.0
 yarl==1.9.2
--- a/examples/langchain-python-rag-privategpt/.gitignore
+++ b/examples/langchain-python-rag-privategpt/.gitignore
@@ -0,0 +1,170 @@
 # OSX
 .DS_STORE
 # Models
 models/
 # Local Chroma db
 .chroma/
 db/
 # Byte-compiled / optimized / DLL files
 __pycache__/
 *.py[cod]
 *$py.class
 # C extensions
 *.so
 # Distribution / packaging
 .Python
 build/
 develop-eggs/
 dist/
 downloads/
 eggs/
 .eggs/
 lib/
 lib64/
 parts/
 sdist/
 var/
 wheels/
 share/python-wheels/
 *.egg-info/
 .installed.cfg
 *.egg
 MANIFEST
 # PyInstaller
 #  Usually these files are written by a python script from a template
 #  before PyInstaller builds the exe, so as to inject date/other infos into it.
 *.manifest
 *.spec
 # Installer logs
 pip-log.txt
 pip-delete-this-directory.txt
 # Unit test / coverage reports
 htmlcov/
 .tox/
 .nox/
 .coverage
 .coverage.*
 .cache
 nosetests.xml
 coverage.xml
 *.cover
 *.py,cover
 .hypothesis/
 .pytest_cache/
 cover/
 # Translations
 *.mo
 *.pot
 # Django stuff:
 *.log
 local_settings.py
 db.sqlite3
 db.sqlite3-journal
 # Flask stuff:
 instance/
 .webassets-cache
 # Scrapy stuff:
 .scrapy
 # Sphinx documentation
 docs/_build/
 # PyBuilder
 .pybuilder/
 target/
 # Jupyter Notebook
 .ipynb_checkpoints
 # IPython
 profile_default/
 ipython_config.py
 # pyenv
 #   For a library or package, you might want to ignore these files since the code is
 #   intended to run in multiple environments; otherwise, check them in:
 # .python-version
 # pipenv
 #   According to pypa/pipenv#598, it is recommended to include Pipfile.lock in version control.
 #   However, in case of collaboration, if having platform-specific dependencies or dependencies
 #   having no cross-platform support, pipenv may install dependencies that don't work, or not
 #   install all needed dependencies.
 #Pipfile.lock
 # poetry
 #   Similar to Pipfile.lock, it is generally recommended to include poetry.lock in version control.
 #   This is especially recommended for binary packages to ensure reproducibility, and is more
 #   commonly ignored for libraries.
 #   https://python-poetry.org/docs/basic-usage/#commit-your-poetrylock-file-to-version-control
 #poetry.lock
 # pdm
 #   Similar to Pipfile.lock, it is generally recommended to include pdm.lock in version control.
 #pdm.lock
 #   pdm stores project-wide configurations in .pdm.toml, but it is recommended to not include it
 #   in version control.
 #   https://pdm.fming.dev/#use-with-ide
 .pdm.toml
 # PEP 582; used by e.g. github.com/David-OConnor/pyflow and github.com/pdm-project/pdm
 __pypackages__/
 # Celery stuff
 celerybeat-schedule
 celerybeat.pid
 # SageMath parsed files
 *.sage.py
 # Environments
 .env
 .venv
 env/
 venv/
 ENV/
 env.bak/
 venv.bak/
 # Spyder project settings
 .spyderproject
 .spyproject
 # Rope project settings
 .ropeproject
 # mkdocs documentation
 /site
 # mypy
 .mypy_cache/
 .dmypy.json
 dmypy.json
 # Pyre type checker
 .pyre/
 # pytype static type analyzer
 .pytype/
 # Cython debug symbols
 cython_debug/
 # PyCharm
 #  JetBrains specific template is maintained in a separate JetBrains.gitignore that can
 #  be found at https://github.com/github/gitignore/blob/main/Global/JetBrains.gitignore
 #  and can be added to the global gitignore or merged into this file.  For a more nuclear
 #  option (not recommended) you can uncomment the following to ignore the entire idea folder.
 #.idea/
--- a/examples/langchain-python-rag-privategpt/LICENSE
+++ b/examples/langchain-python-rag-privategpt/LICENSE
@@ -0,0 +1,201 @@
                                 Apache License
                           Version 2.0, January 2004
                        http://www.apache.org/licenses/
   TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
   1. Definitions.
      "License" shall mean the terms and conditions for use, reproduction,
      and distribution as defined by Sections 1 through 9 of this document.
      "Licensor" shall mean the copyright owner or entity authorized by
      the copyright owner that is granting the License.
      "Legal Entity" shall mean the union of the acting entity and all
      other entities that control, are controlled by, or are under common
      control with that entity. For the purposes of this definition,
      "control" means (i) the power, direct or indirect, to cause the
      direction or management of such entity, whether by contract or
      otherwise, or (ii) ownership of fifty percent (50%) or more of the
      outstanding shares, or (iii) beneficial ownership of such entity.
      "You" (or "Your") shall mean an individual or Legal Entity
      exercising permissions granted by this License.
      "Source" form shall mean the preferred form for making modifications,
      including but not limited to software source code, documentation
      source, and configuration files.
      "Object" form shall mean any form resulting from mechanical
      transformation or translation of a Source form, including but
      not limited to compiled object code, generated documentation,
      and conversions to other media types.
      "Work" shall mean the work of authorship, whether in Source or
      Object form, made available under the License, as indicated by a
      copyright notice that is included in or attached to the work
      (an example is provided in the Appendix below).
      "Derivative Works" shall mean any work, whether in Source or Object
      form, that is based on (or derived from) the Work and for which the
      editorial revisions, annotations, elaborations, or other modifications
      represent, as a whole, an original work of authorship. For the purposes
      of this License, Derivative Works shall not include works that remain
      separable from, or merely link (or bind by name) to the interfaces of,
      the Work and Derivative Works thereof.
      "Contribution" shall mean any work of authorship, including
      the original version of the Work and any modifications or additions
      to that Work or Derivative Works thereof, that is intentionally
      submitted to Licensor for inclusion in the Work by the copyright owner
      or by an individual or Legal Entity authorized to submit on behalf of
      the copyright owner. For the purposes of this definition, "submitted"
      means any form of electronic, verbal, or written communication sent
      to the Licensor or its representatives, including but not limited to
      communication on electronic mailing lists, source code control systems,
      and issue tracking systems that are managed by, or on behalf of, the
      Licensor for the purpose of discussing and improving the Work, but
      excluding communication that is conspicuously marked or otherwise
      designated in writing by the copyright owner as "Not a Contribution."
      "Contributor" shall mean Licensor and any individual or Legal Entity
      on behalf of whom a Contribution has been received by Licensor and
      subsequently incorporated within the Work.
   2. Grant of Copyright License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      copyright license to reproduce, prepare Derivative Works of,
      publicly display, publicly perform, sublicense, and distribute the
      Work and such Derivative Works in Source or Object form.
   3. Grant of Patent License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      (except as stated in this section) patent license to make, have made,
      use, offer to sell, sell, import, and otherwise transfer the Work,
      where such license applies only to those patent claims licensable
      by such Contributor that are necessarily infringed by their
      Contribution(s) alone or by combination of their Contribution(s)
      with the Work to which such Contribution(s) was submitted. If You
      institute patent litigation against any entity (including a
      cross-claim or counterclaim in a lawsuit) alleging that the Work
      or a Contribution incorporated within the Work constitutes direct
      or contributory patent infringement, then any patent licenses
      granted to You under this License for that Work shall terminate
      as of the date such litigation is filed.
   4. Redistribution. You may reproduce and distribute copies of the
      Work or Derivative Works thereof in any medium, with or without
      modifications, and in Source or Object form, provided that You
      meet the following conditions:
      (a) You must give any other recipients of the Work or
          Derivative Works a copy of this License; and
      (b) You must cause any modified files to carry prominent notices
          stating that You changed the files; and
      (c) You must retain, in the Source form of any Derivative Works
          that You distribute, all copyright, patent, trademark, and
          attribution notices from the Source form of the Work,
          excluding those notices that do not pertain to any part of
          the Derivative Works; and
      (d) If the Work includes a "NOTICE" text file as part of its
          distribution, then any Derivative Works that You distribute must
          include a readable copy of the attribution notices contained
          within such NOTICE file, excluding those notices that do not
          pertain to any part of the Derivative Works, in at least one
          of the following places: within a NOTICE text file distributed
          as part of the Derivative Works; within the Source form or
          documentation, if provided along with the Derivative Works; or,
          within a display generated by the Derivative Works, if and
          wherever such third-party notices normally appear. The contents
          of the NOTICE file are for informational purposes only and
          do not modify the License. You may add Your own attribution
          notices within Derivative Works that You distribute, alongside
          or as an addendum to the NOTICE text from the Work, provided
          that such additional attribution notices cannot be construed
          as modifying the License.
      You may add Your own copyright statement to Your modifications and
      may provide additional or different license terms and conditions
      for use, reproduction, or distribution of Your modifications, or
      for any such Derivative Works as a whole, provided Your use,
      reproduction, and distribution of the Work otherwise complies with
      the conditions stated in this License.
   5. Submission of Contributions. Unless You explicitly state otherwise,
      any Contribution intentionally submitted for inclusion in the Work
      by You to the Licensor shall be under the terms and conditions of
      this License, without any additional terms or conditions.
      Notwithstanding the above, nothing herein shall supersede or modify
      the terms of any separate license agreement you may have executed
      with Licensor regarding such Contributions.
   6. Trademarks. This License does not grant permission to use the trade
      names, trademarks, service marks, or product names of the Licensor,
      except as required for reasonable and customary use in describing the
      origin of the Work and reproducing the content of the NOTICE file.
   7. Disclaimer of Warranty. Unless required by applicable law or
      agreed to in writing, Licensor provides the Work (and each
      Contributor provides its Contributions) on an "AS IS" BASIS,
      WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
      implied, including, without limitation, any warranties or conditions
      of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
      PARTICULAR PURPOSE. You are solely responsible for determining the
      appropriateness of using or redistributing the Work and assume any
      risks associated with Your exercise of permissions under this License.
   8. Limitation of Liability. In no event and under no legal theory,
      whether in tort (including negligence), contract, or otherwise,
      unless required by applicable law (such as deliberate and grossly
      negligent acts) or agreed to in writing, shall any Contributor be
      liable to You for damages, including any direct, indirect, special,
      incidental, or consequential damages of any character arising as a
      result of this License or out of the use or inability to use the
      Work (including but not limited to damages for loss of goodwill,
      work stoppage, computer failure or malfunction, or any and all
      other commercial damages or losses), even if such Contributor
      has been advised of the possibility of such damages.
   9. Accepting Warranty or Additional Liability. While redistributing
      the Work or Derivative Works thereof, You may choose to offer,
      and charge a fee for, acceptance of support, warranty, indemnity,
      or other liability obligations and/or rights consistent with this
      License. However, in accepting such obligations, You may act only
      on Your own behalf and on Your sole responsibility, not on behalf
      of any other Contributor, and only if You agree to indemnify,
      defend, and hold each Contributor harmless for any liability
      incurred by, or claims asserted against, such Contributor by reason
      of your accepting any such warranty or additional liability.
   END OF TERMS AND CONDITIONS
   APPENDIX: How to apply the Apache License to your work.
      To apply the Apache License to your work, attach the following
      boilerplate notice, with the fields enclosed by brackets "[]"
      replaced with your own identifying information. (Don't include
      the brackets!)  The text should be enclosed in the appropriate
      comment syntax for the file format. We also recommend that a
      file or class name and description of purpose be included on the
      same "printed page" as the copyright notice for easier
      identification within third-party archives.
   Copyright [yyyy] [name of copyright owner]
   Licensed under the Apache License, Version 2.0 (the "License");
   you may not use this file except in compliance with the License.
   You may obtain a copy of the License at
       http://www.apache.org/licenses/LICENSE-2.0
   Unless required by applicable law or agreed to in writing, software
   distributed under the License is distributed on an "AS IS" BASIS,
   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
   See the License for the specific language governing permissions and
   limitations under the License.
--- a/examples/langchain-python-rag-privategpt/README.md
+++ b/examples/langchain-python-rag-privategpt/README.md
@@ -0,0 +1,91 @@
 # PrivateGPT with Llama 2 uncensored
 https://github.com/jmorganca/ollama/assets/3325447/20cf8ec6-ff25-42c6-bdd8-9be594e3ce1b
 > Note: this example is a slightly modified version of PrivateGPT using models such as Llama 2 Uncensored. All credit for PrivateGPT goes to Iván Martínez who is the creator of it, and you can find his GitHub repo [here](https://github.com/imartinez/privateGPT).
 ### Setup
 Set up a virtual environment (optional):
 ```
 python3 -m venv .venv
 source .venv/bin/activate
 ```
 Install the Python dependencies:
 ```shell
 pip install -r requirements.txt
 ```
 Pull the model you'd like to use:
 ```
 ollama pull llama2-uncensored
 ```
 ### Getting WeWork's latest quarterly earnings report (10-Q)
 ```
 mkdir source_documents
 curl https://d18rn0p25nwr6d.cloudfront.net/CIK-0001813756/975b3e9b-268e-4798-a9e4-2a9a7c92dc10.pdf -o source_documents/wework.pdf
 ```
 ### Ingesting files
 ```shell
 python ingest.py
 ```
 Output should look like this:
 ```shell
 Creating new vectorstore
 Loading documents from source_documents
 Loading new documents: 100%|██████████████████████| 1/1 [00:01<00:00,  1.73s/it]
 Loaded 1 new documents from source_documents
 Split into 90 chunks of text (max. 500 tokens each)
 Creating embeddings. May take some minutes...
 Using embedded DuckDB with persistence: data will be stored in: db
 Ingestion complete! You can now run privateGPT.py to query your documents
 ```
 ### Ask questions
 ```shell
 python privateGPT.py
 Enter a query: How many locations does WeWork have?
 > Answer (took 17.7 s.):
 As of June 2023, WeWork has 777 locations worldwide, including 610 Consolidated Locations (as defined in the section entitled Key Performance Indicators).
 ```
 ### Try a different model:
 ```
 ollama pull llama2:13b
 MODEL=llama2:13b python privateGPT.py
 ```
 ## Adding more files
 Put any and all your files into the `source_documents` directory
 The supported extensions are:
 - `.csv`: CSV,
 - `.docx`: Word Document,
 - `.doc`: Word Document,
 - `.enex`: EverNote,
 - `.eml`: Email,
 - `.epub`: EPub,
 - `.html`: HTML File,
 - `.md`: Markdown,
 - `.msg`: Outlook Message,
 - `.odt`: Open Document Text,
 - `.pdf`: Portable Document Format (PDF),
 - `.pptx` : PowerPoint Document,
 - `.ppt` : PowerPoint Document,
 - `.txt`: Text file (UTF-8),
--- a/examples/langchain-python-rag-privategpt/constants.py
+++ b/examples/langchain-python-rag-privategpt/constants.py
@@ -0,0 +1,12 @@
 import os
 from chromadb.config import Settings
 # Define the folder for storing database
 PERSIST_DIRECTORY = os.environ.get('PERSIST_DIRECTORY', 'db')
 # Define the Chroma settings
 CHROMA_SETTINGS = Settings(
        chroma_db_impl='duckdb+parquet',
        persist_directory=PERSIST_DIRECTORY,
        anonymized_telemetry=False
 )
--- a/examples/langchain-python-rag-privategpt/ingest.py
+++ b/examples/langchain-python-rag-privategpt/ingest.py
@@ -0,0 +1,161 @@
 #!/usr/bin/env python3
 import os
 import glob
 from typing import List
 from multiprocessing import Pool
 from tqdm import tqdm
 from langchain.document_loaders import (
    CSVLoader,
    EverNoteLoader,
    PyMuPDFLoader,
    TextLoader,
    UnstructuredEmailLoader,
    UnstructuredEPubLoader,
    UnstructuredHTMLLoader,
    UnstructuredMarkdownLoader,
    UnstructuredODTLoader,
    UnstructuredPowerPointLoader,
    UnstructuredWordDocumentLoader,
 )
 from langchain.text_splitter import RecursiveCharacterTextSplitter
 from langchain.vectorstores import Chroma
 from langchain.embeddings import HuggingFaceEmbeddings
 from langchain.docstore.document import Document
 from constants import CHROMA_SETTINGS
 # Load environment variables
 persist_directory = os.environ.get('PERSIST_DIRECTORY', 'db')
 source_directory = os.environ.get('SOURCE_DIRECTORY', 'source_documents')
 embeddings_model_name = os.environ.get('EMBEDDINGS_MODEL_NAME', 'all-MiniLM-L6-v2')
 chunk_size = 500
 chunk_overlap = 50
 # Custom document loaders
 class MyElmLoader(UnstructuredEmailLoader):
    """Wrapper to fallback to text/plain when default does not work"""
    def load(self) -> List[Document]:
        """Wrapper adding fallback for elm without html"""
        try:
            try:
                doc = UnstructuredEmailLoader.load(self)
            except ValueError as e:
                if 'text/html content not found in email' in str(e):
                    # Try plain text
                    self.unstructured_kwargs["content_source"]="text/plain"
                    doc = UnstructuredEmailLoader.load(self)
                else:
                    raise
        except Exception as e:
            # Add file_path to exception message
            raise type(e)(f"{self.file_path}: {e}") from e
        return doc
 # Map file extensions to document loaders and their arguments
 LOADER_MAPPING = {
    ".csv": (CSVLoader, {}),
    # ".docx": (Docx2txtLoader, {}),
    ".doc": (UnstructuredWordDocumentLoader, {}),
    ".docx": (UnstructuredWordDocumentLoader, {}),
    ".enex": (EverNoteLoader, {}),
    ".eml": (MyElmLoader, {}),
    ".epub": (UnstructuredEPubLoader, {}),
    ".html": (UnstructuredHTMLLoader, {}),
    ".md": (UnstructuredMarkdownLoader, {}),
    ".odt": (UnstructuredODTLoader, {}),
    ".pdf": (PyMuPDFLoader, {}),
    ".ppt": (UnstructuredPowerPointLoader, {}),
    ".pptx": (UnstructuredPowerPointLoader, {}),
    ".txt": (TextLoader, {"encoding": "utf8"}),
    # Add more mappings for other file extensions and loaders as needed
 }
 def load_single_document(file_path: str) -> List[Document]:
    ext = "." + file_path.rsplit(".", 1)[-1]
    if ext in LOADER_MAPPING:
        loader_class, loader_args = LOADER_MAPPING[ext]
        loader = loader_class(file_path, **loader_args)
        return loader.load()
    raise ValueError(f"Unsupported file extension '{ext}'")
 def load_documents(source_dir: str, ignored_files: List[str] = []) -> List[Document]:
    """
    Loads all documents from the source documents directory, ignoring specified files
    """
    all_files = []
    for ext in LOADER_MAPPING:
        all_files.extend(
            glob.glob(os.path.join(source_dir, f"**/*{ext}"), recursive=True)
        )
    filtered_files = [file_path for file_path in all_files if file_path not in ignored_files]
    with Pool(processes=os.cpu_count()) as pool:
        results = []
        with tqdm(total=len(filtered_files), desc='Loading new documents', ncols=80) as pbar:
            for i, docs in enumerate(pool.imap_unordered(load_single_document, filtered_files)):
                results.extend(docs)
                pbar.update()
    return results
 def process_documents(ignored_files: List[str] = []) -> List[Document]:
    """
    Load documents and split in chunks
    """
    print(f"Loading documents from {source_directory}")
    documents = load_documents(source_directory, ignored_files)
    if not documents:
        print("No new documents to load")
        exit(0)
    print(f"Loaded {len(documents)} new documents from {source_directory}")
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=chunk_size, chunk_overlap=chunk_overlap)
    texts = text_splitter.split_documents(documents)
    print(f"Split into {len(texts)} chunks of text (max. {chunk_size} tokens each)")
    return texts
 def does_vectorstore_exist(persist_directory: str) -> bool:
    """
    Checks if vectorstore exists
    """
    if os.path.exists(os.path.join(persist_directory, 'index')):
        if os.path.exists(os.path.join(persist_directory, 'chroma-collections.parquet')) and os.path.exists(os.path.join(persist_directory, 'chroma-embeddings.parquet')):
            list_index_files = glob.glob(os.path.join(persist_directory, 'index/*.bin'))
            list_index_files += glob.glob(os.path.join(persist_directory, 'index/*.pkl'))
            # At least 3 documents are needed in a working vectorstore
            if len(list_index_files) > 3:
                return True
    return False
 def main():
    # Create embeddings
    embeddings = HuggingFaceEmbeddings(model_name=embeddings_model_name)
    if does_vectorstore_exist(persist_directory):
        # Update and store locally vectorstore
        print(f"Appending to existing vectorstore at {persist_directory}")
        db = Chroma(persist_directory=persist_directory, embedding_function=embeddings, client_settings=CHROMA_SETTINGS)
        collection = db.get()
        texts = process_documents([metadata['source'] for metadata in collection['metadatas']])
        print(f"Creating embeddings. May take some minutes...")
        db.add_documents(texts)
    else:
        # Create and store locally vectorstore
        print("Creating new vectorstore")
        texts = process_documents()
        print(f"Creating embeddings. May take some minutes...")
        db = Chroma.from_documents(texts, embeddings, persist_directory=persist_directory, client_settings=CHROMA_SETTINGS)
    db.persist()
    db = None
    print(f"Ingestion complete! You can now run privateGPT.py to query your documents")
 if __name__ == "__main__":
    main()
--- a/examples/langchain-python-rag-privategpt/poetry.lock
+++ b/examples/langchain-python-rag-privategpt/poetry.lock
--- a/examples/langchain-python-rag-privategpt/privateGPT.py
+++ b/examples/langchain-python-rag-privategpt/privateGPT.py
@@ -0,0 +1,71 @@
 #!/usr/bin/env python3
 from langchain.chains import RetrievalQA
 from langchain.embeddings import HuggingFaceEmbeddings
 from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler
 from langchain.vectorstores import Chroma
 from langchain.llms import Ollama
 import os
 import argparse
 import time
 model = os.environ.get("MODEL", "llama2-uncensored")
 # For embeddings model, the example uses a sentence-transformers model
 # https://www.sbert.net/docs/pretrained_models.html 
 # "The all-mpnet-base-v2 model provides the best quality, while all-MiniLM-L6-v2 is 5 times faster and still offers good quality."
 embeddings_model_name = os.environ.get("EMBEDDINGS_MODEL_NAME", "all-MiniLM-L6-v2")
 persist_directory = os.environ.get("PERSIST_DIRECTORY", "db")
 target_source_chunks = int(os.environ.get('TARGET_SOURCE_CHUNKS',4))
 from constants import CHROMA_SETTINGS
 def main():
    # Parse the command line arguments
    args = parse_arguments()
    embeddings = HuggingFaceEmbeddings(model_name=embeddings_model_name)
    db = Chroma(persist_directory=persist_directory, embedding_function=embeddings, client_settings=CHROMA_SETTINGS)
    retriever = db.as_retriever(search_kwargs={"k": target_source_chunks})
    # activate/deactivate the streaming StdOut callback for LLMs
    callbacks = [] if args.mute_stream else [StreamingStdOutCallbackHandler()]
    llm = Ollama(model=model, callbacks=callbacks)
    qa = RetrievalQA.from_chain_type(llm=llm, chain_type="stuff", retriever=retriever, return_source_documents= not args.hide_source)
    # Interactive questions and answers
    while True:
        query = input("\nEnter a query: ")
        if query == "exit":
            break
        if query.strip() == "":
            continue
        # Get the answer from the chain
        start = time.time()
        res = qa(query)
        answer, docs = res['result'], [] if args.hide_source else res['source_documents']
        end = time.time()
        # Print the result
        print("\n\n> Question:")
        print(query)
        print(answer)
        # Print the relevant sources used for the answer
        for document in docs:
            print("\n> " + document.metadata["source"] + ":")
            print(document.page_content)
 def parse_arguments():
    parser = argparse.ArgumentParser(description='privateGPT: Ask questions to your documents without an internet connection, '
                                                 'using the power of LLMs.')
    parser.add_argument("--hide-source", "-S", action='store_true',
                        help='Use this flag to disable printing of source documents used for answers.')
    parser.add_argument("--mute-stream", "-M",
                        action='store_true',
                        help='Use this flag to disable the streaming StdOut callback for LLMs.')
    return parser.parse_args()
 if __name__ == "__main__":
    main()
--- a/examples/langchain-python-rag-privategpt/pyproject.toml
+++ b/examples/langchain-python-rag-privategpt/pyproject.toml
@@ -0,0 +1,26 @@
 [tool.poetry]
 name = "privategpt"
 version = "0.1.0"
 description = ""
 authors = ["Ivan Martinez <ivanmartit@gmail.com>"]
 license = "Apache Version 2.0"
 readme = "README.md"
 [tool.poetry.dependencies]
 python = "^3.10"
 langchain = "0.0.261"
 gpt4all = "^1.0.3"
 chromadb = "^0.3.26"
 PyMuPDF = "^1.22.5"
 python-dotenv = "^1.0.0"
 unstructured = "^0.8.0"
 extract-msg = "^0.41.5"
 tabulate = "^0.9.0"
 pandoc = "^2.3"
 pypandoc = "^1.11"
 tqdm = "^4.65.0"
 sentence-transformers = "^2.2.2"
 [build-system]
 requires = ["poetry-core"]
 build-backend = "poetry.core.masonry.api"
--- a/examples/langchain-python-rag-privategpt/requirements.txt
+++ b/examples/langchain-python-rag-privategpt/requirements.txt
--- a/examples/langchain-python-rag-websummary/README.md
+++ b/examples/langchain-python-rag-websummary/README.md
@@ -0,0 +1,15 @@
 # LangChain Web Summarization
 This example summarizes a website
 ## Setup
 ```
 pip install -r requirements.txt
 ```
 ## Run
 ```
 python main.py
 ```
--- a/examples/langchain-python-rag-websummary/main.py
+++ b/examples/langchain-python-rag-websummary/main.py
@@ -0,0 +1,12 @@
 from langchain.llms import Ollama
 from langchain.document_loaders import WebBaseLoader
 from langchain.chains.summarize import load_summarize_chain
 loader = WebBaseLoader("https://ollama.ai/blog/run-llama2-uncensored-locally")
 docs = loader.load()
 llm = Ollama(model="llama2")
 chain = load_summarize_chain(llm, chain_type="stuff")
 result = chain.run(docs)
 print(result)
--- a/examples/langchain-python-rag-websummary/requirements.txt
+++ b/examples/langchain-python-rag-websummary/requirements.txt
@@ -0,0 +1,2 @@
 langchain==0.0.259
 bs4==0.0.1
--- a/examples/langchain-python-simple/README.md
+++ b/examples/langchain-python-simple/README.md
@@ -0,0 +1,21 @@
 # LangChain
 This example is a basic "hello world" of using LangChain with Ollama.
 ## Setup
 ```
 pip install -r requirements.txt
 ```
 ## Run
 ```
 python main.py
 ```
 Running this example will print the response for "hello":
 ```
 Hello! It's nice to meet you. hopefully you are having a great day! Is there something I can help you with or would you like to chat?
 ```
--- a/examples/langchain-python-simple/main.py
+++ b/examples/langchain-python-simple/main.py
@@ -0,0 +1,4 @@
 from langchain.llms import Ollama
 llm = Ollama(model="llama2")
 res = llm.predict("hello")
 print (res)
--- a/examples/langchain-python-simple/requirements.txt
+++ b/examples/langchain-python-simple/requirements.txt
@@ -0,0 +1 @@
 langchain==0.0.259
--- a/examples/langchain-typescript-simple/README.md
+++ b/examples/langchain-typescript-simple/README.md
@@ -0,0 +1,21 @@
 # LangChain
 This example is a basic "hello world" of using LangChain with Ollama using Node.js and Typescript.
 ## Setup
 ```shell
 npm install
 ```
 ## Run
 ```shell
 ts-node main.ts
 ```
 Running this example will print the response for "hello":
 ```plaintext
 Hello! It's nice to meet you. hopefully you are having a great day! Is there something I can help you with or would you like to chat?
 ```
--- a/examples/langchain-typescript-simple/main.ts
+++ b/examples/langchain-typescript-simple/main.ts
@@ -0,0 +1,15 @@
 import { Ollama} from 'langchain/llms/ollama';
 async function main() {
  const ollama = new Ollama({
    model: 'mistral'    
    // other parameters can be found at https://js.langchain.com/docs/api/llms_ollama/classes/Ollama
  })
  const stream = await ollama.stream("Hello");
  for await (const chunk of stream) {
    process.stdout.write(chunk);
  }
 }
 main();
--- a/examples/langchain-typescript-simple/package-lock.json
+++ b/examples/langchain-typescript-simple/package-lock.json
@@ -0,0 +1,997 @@
 {
  "name": "with-langchain-typescript-simplegenerate",
  "lockfileVersion": 3,
  "requires": true,
  "packages": {
    "": {
      "dependencies": {
        "langchain": "^0.0.165"
      },
      "devDependencies": {
        "typescript": "^5.2.2"
      }
    },
    "node_modules/@anthropic-ai/sdk": {
      "version": "0.6.2",
      "resolved": "https://registry.npmjs.org/@anthropic-ai/sdk/-/sdk-0.6.2.tgz",
      "integrity": "sha512-fB9PUj9RFT+XjkL+E9Ol864ZIJi+1P8WnbHspN3N3/GK2uSzjd0cbVIKTGgf4v3N8MwaQu+UWnU7C4BG/fap/g==",
      "dependencies": {
        "@types/node": "^18.11.18",
        "@types/node-fetch": "^2.6.4",
        "abort-controller": "^3.0.0",
        "agentkeepalive": "^4.2.1",
        "digest-fetch": "^1.3.0",
        "form-data-encoder": "1.7.2",
        "formdata-node": "^4.3.2",
        "node-fetch": "^2.6.7"
      }
    },
    "node_modules/@types/node": {
      "version": "18.18.4",
      "resolved": "https://registry.npmjs.org/@types/node/-/node-18.18.4.tgz",
      "integrity": "sha512-t3rNFBgJRugIhackit2mVcLfF6IRc0JE4oeizPQL8Zrm8n2WY/0wOdpOPhdtG0V9Q2TlW/axbF1MJ6z+Yj/kKQ=="
    },
    "node_modules/@types/node-fetch": {
      "version": "2.6.6",
      "resolved": "https://registry.npmjs.org/@types/node-fetch/-/node-fetch-2.6.6.tgz",
      "integrity": "sha512-95X8guJYhfqiuVVhRFxVQcf4hW/2bCuoPwDasMf/531STFoNoWTT7YDnWdXHEZKqAGUigmpG31r2FE70LwnzJw==",
      "dependencies": {
        "@types/node": "*",
        "form-data": "^4.0.0"
      }
    },
    "node_modules/@types/retry": {
      "version": "0.12.0",
      "resolved": "https://registry.npmjs.org/@types/retry/-/retry-0.12.0.tgz",
      "integrity": "sha512-wWKOClTTiizcZhXnPY4wikVAwmdYHp8q6DmC+EJUzAMsycb7HB32Kh9RN4+0gExjmPmZSAQjgURXIGATPegAvA=="
    },
    "node_modules/@types/uuid": {
      "version": "9.0.5",
      "resolved": "https://registry.npmjs.org/@types/uuid/-/uuid-9.0.5.tgz",
      "integrity": "sha512-xfHdwa1FMJ082prjSJpoEI57GZITiQz10r3vEJCHa2khEFQjKy91aWKz6+zybzssCvXUwE1LQWgWVwZ4nYUvHQ=="
    },
    "node_modules/abort-controller": {
      "version": "3.0.0",
      "resolved": "https://registry.npmjs.org/abort-controller/-/abort-controller-3.0.0.tgz",
      "integrity": "sha512-h8lQ8tacZYnR3vNQTgibj+tODHI5/+l06Au2Pcriv/Gmet0eaj4TwWH41sO9wnHDiQsEj19q0drzdWdeAHtweg==",
      "dependencies": {
        "event-target-shim": "^5.0.0"
      },
      "engines": {
        "node": ">=6.5"
      }
    },
    "node_modules/agentkeepalive": {
      "version": "4.5.0",
      "resolved": "https://registry.npmjs.org/agentkeepalive/-/agentkeepalive-4.5.0.tgz",
      "integrity": "sha512-5GG/5IbQQpC9FpkRGsSvZI5QYeSCzlJHdpBQntCsuTOxhKD8lqKhrleg2Yi7yvMIf82Ycmmqln9U8V9qwEiJew==",
      "dependencies": {
        "humanize-ms": "^1.2.1"
      },
      "engines": {
        "node": ">= 8.0.0"
      }
    },
    "node_modules/ansi-styles": {
      "version": "5.2.0",
      "resolved": "https://registry.npmjs.org/ansi-styles/-/ansi-styles-5.2.0.tgz",
      "integrity": "sha512-Cxwpt2SfTzTtXcfOlzGEee8O+c+MmUgGrNiBcXnuWxuFJHe6a5Hz7qwhwe5OgaSYI0IJvkLqWX1ASG+cJOkEiA==",
      "engines": {
        "node": ">=10"
      },
      "funding": {
        "url": "https://github.com/chalk/ansi-styles?sponsor=1"
      }
    },
    "node_modules/argparse": {
      "version": "2.0.1",
      "resolved": "https://registry.npmjs.org/argparse/-/argparse-2.0.1.tgz",
      "integrity": "sha512-8+9WqebbFzpX9OR+Wa6O29asIogeRMzcGtAINdpMHHyAg10f05aSFVBbcEqGf/PXw1EjAZ+q2/bEBg3DvurK3Q=="
    },
    "node_modules/asynckit": {
      "version": "0.4.0",
      "resolved": "https://registry.npmjs.org/asynckit/-/asynckit-0.4.0.tgz",
      "integrity": "sha512-Oei9OH4tRh0YqU3GxhX79dM/mwVgvbZJaSNaRk+bshkj0S5cfHcgYakreBjrHwatXKbz+IoIdYLxrKim2MjW0Q=="
    },
    "node_modules/base-64": {
      "version": "0.1.0",
      "resolved": "https://registry.npmjs.org/base-64/-/base-64-0.1.0.tgz",
      "integrity": "sha512-Y5gU45svrR5tI2Vt/X9GPd3L0HNIKzGu202EjxrXMpuc2V2CiKgemAbUUsqYmZJvPtCXoUKjNZwBJzsNScUbXA=="
    },
    "node_modules/base64-js": {
      "version": "1.5.1",
      "resolved": "https://registry.npmjs.org/base64-js/-/base64-js-1.5.1.tgz",
      "integrity": "sha512-AKpaYlHn8t4SVbOHCy+b5+KKgvR4vrsD8vbvrbiQJps7fKDTkjkDry6ji0rUJjC0kzbNePLwzxq8iypo41qeWA==",
      "funding": [
        {
          "type": "github",
          "url": "https://github.com/sponsors/feross"
        },
        {
          "type": "patreon",
          "url": "https://www.patreon.com/feross"
        },
        {
          "type": "consulting",
          "url": "https://feross.org/support"
        }
      ]
    },
    "node_modules/binary-extensions": {
      "version": "2.2.0",
      "resolved": "https://registry.npmjs.org/binary-extensions/-/binary-extensions-2.2.0.tgz",
      "integrity": "sha512-jDctJ/IVQbZoJykoeHbhXpOlNBqGNcwXJKJog42E5HDPUwQTSdjCHdihjj0DlnheQ7blbT6dHOafNAiS8ooQKA==",
      "engines": {
        "node": ">=8"
      }
    },
    "node_modules/binary-search": {
      "version": "1.3.6",
      "resolved": "https://registry.npmjs.org/binary-search/-/binary-search-1.3.6.tgz",
      "integrity": "sha512-nbE1WxOTTrUWIfsfZ4aHGYu5DOuNkbxGokjV6Z2kxfJK3uaAb8zNK1muzOeipoLHZjInT4Br88BHpzevc681xA=="
    },
    "node_modules/camelcase": {
      "version": "6.3.0",
      "resolved": "https://registry.npmjs.org/camelcase/-/camelcase-6.3.0.tgz",
      "integrity": "sha512-Gmy6FhYlCY7uOElZUSbxo2UCDH8owEk996gkbrpsgGtrJLM3J7jGxl9Ic7Qwwj4ivOE5AWZWRMecDdF7hqGjFA==",
      "engines": {
        "node": ">=10"
      },
      "funding": {
        "url": "https://github.com/sponsors/sindresorhus"
      }
    },
    "node_modules/charenc": {
      "version": "0.0.2",
      "resolved": "https://registry.npmjs.org/charenc/-/charenc-0.0.2.tgz",
      "integrity": "sha512-yrLQ/yVUFXkzg7EDQsPieE/53+0RlaWTs+wBrvW36cyilJ2SaDWfl4Yj7MtLTXleV9uEKefbAGUPv2/iWSooRA==",
      "engines": {
        "node": "*"
      }
    },
    "node_modules/combined-stream": {
      "version": "1.0.8",
      "resolved": "https://registry.npmjs.org/combined-stream/-/combined-stream-1.0.8.tgz",
      "integrity": "sha512-FQN4MRfuJeHf7cBbBMJFXhKSDq+2kAArBlmRBvcvFE5BB1HZKXtSFASDhdlz9zOYwxh8lDdnvmMOe/+5cdoEdg==",
      "dependencies": {
        "delayed-stream": "~1.0.0"
      },
      "engines": {
        "node": ">= 0.8"
      }
    },
    "node_modules/commander": {
      "version": "10.0.1",
      "resolved": "https://registry.npmjs.org/commander/-/commander-10.0.1.tgz",
      "integrity": "sha512-y4Mg2tXshplEbSGzx7amzPwKKOCGuoSRP/CjEdwwk0FOGlUbq6lKuoyDZTNZkmxHdJtp54hdfY/JUrdL7Xfdug==",
      "engines": {
        "node": ">=14"
      }
    },
    "node_modules/crypt": {
      "version": "0.0.2",
      "resolved": "https://registry.npmjs.org/crypt/-/crypt-0.0.2.tgz",
      "integrity": "sha512-mCxBlsHFYh9C+HVpiEacem8FEBnMXgU9gy4zmNC+SXAZNB/1idgp/aulFJ4FgCi7GPEVbfyng092GqL2k2rmow==",
      "engines": {
        "node": "*"
      }
    },
    "node_modules/decamelize": {
      "version": "1.2.0",
      "resolved": "https://registry.npmjs.org/decamelize/-/decamelize-1.2.0.tgz",
      "integrity": "sha512-z2S+W9X73hAUUki+N+9Za2lBlun89zigOyGrsax+KUQ6wKW4ZoWpEYBkGhQjwAjjDCkWxhY0VKEhk8wzY7F5cA==",
      "engines": {
        "node": ">=0.10.0"
      }
    },
    "node_modules/delayed-stream": {
      "version": "1.0.0",
      "resolved": "https://registry.npmjs.org/delayed-stream/-/delayed-stream-1.0.0.tgz",
      "integrity": "sha512-ZySD7Nf91aLB0RxL4KGrKHBXl7Eds1DAmEdcoVawXnLD7SDhpNgtuII2aAkg7a7QS41jxPSZ17p4VdGnMHk3MQ==",
      "engines": {
        "node": ">=0.4.0"
      }
    },
    "node_modules/digest-fetch": {
      "version": "1.3.0",
      "resolved": "https://registry.npmjs.org/digest-fetch/-/digest-fetch-1.3.0.tgz",
      "integrity": "sha512-CGJuv6iKNM7QyZlM2T3sPAdZWd/p9zQiRNS9G+9COUCwzWFTs0Xp8NF5iePx7wtvhDykReiRRrSeNb4oMmB8lA==",
      "dependencies": {
        "base-64": "^0.1.0",
        "md5": "^2.3.0"
      }
    },
    "node_modules/event-target-shim": {
      "version": "5.0.1",
      "resolved": "https://registry.npmjs.org/event-target-shim/-/event-target-shim-5.0.1.tgz",
      "integrity": "sha512-i/2XbnSz/uxRCU6+NdVJgKWDTM427+MqYbkQzD321DuCQJUqOuJKIA0IM2+W2xtYHdKOmZ4dR6fExsd4SXL+WQ==",
      "engines": {
        "node": ">=6"
      }
    },
    "node_modules/eventemitter3": {
      "version": "4.0.7",
      "resolved": "https://registry.npmjs.org/eventemitter3/-/eventemitter3-4.0.7.tgz",
      "integrity": "sha512-8guHBZCwKnFhYdHr2ysuRWErTwhoN2X8XELRlrRwpmfeY2jjuUN4taQMsULKUVo1K4DvZl+0pgfyoysHxvmvEw=="
    },
    "node_modules/expr-eval": {
      "version": "2.0.2",
      "resolved": "https://registry.npmjs.org/expr-eval/-/expr-eval-2.0.2.tgz",
      "integrity": "sha512-4EMSHGOPSwAfBiibw3ndnP0AvjDWLsMvGOvWEZ2F96IGk0bIVdjQisOHxReSkE13mHcfbuCiXw+G4y0zv6N8Eg=="
    },
    "node_modules/flat": {
      "version": "5.0.2",
      "resolved": "https://registry.npmjs.org/flat/-/flat-5.0.2.tgz",
      "integrity": "sha512-b6suED+5/3rTpUBdG1gupIl8MPFCAMA0QXwmljLhvCUKcUvdE4gWky9zpuGCcXHOsz4J9wPGNWq6OKpmIzz3hQ==",
      "bin": {
        "flat": "cli.js"
      }
    },
    "node_modules/form-data": {
      "version": "4.0.0",
      "resolved": "https://registry.npmjs.org/form-data/-/form-data-4.0.0.tgz",
      "integrity": "sha512-ETEklSGi5t0QMZuiXoA/Q6vcnxcLQP5vdugSpuAyi6SVGi2clPPp+xgEhuMaHC+zGgn31Kd235W35f7Hykkaww==",
      "dependencies": {
        "asynckit": "^0.4.0",
        "combined-stream": "^1.0.8",
        "mime-types": "^2.1.12"
      },
      "engines": {
        "node": ">= 6"
      }
    },
    "node_modules/form-data-encoder": {
      "version": "1.7.2",
      "resolved": "https://registry.npmjs.org/form-data-encoder/-/form-data-encoder-1.7.2.tgz",
      "integrity": "sha512-qfqtYan3rxrnCk1VYaA4H+Ms9xdpPqvLZa6xmMgFvhO32x7/3J/ExcTd6qpxM0vH2GdMI+poehyBZvqfMTto8A=="
    },
    "node_modules/formdata-node": {
      "version": "4.4.1",
      "resolved": "https://registry.npmjs.org/formdata-node/-/formdata-node-4.4.1.tgz",
      "integrity": "sha512-0iirZp3uVDjVGt9p49aTaqjk84TrglENEDuqfdlZQ1roC9CWlPk6Avf8EEnZNcAqPonwkG35x4n3ww/1THYAeQ==",
      "dependencies": {
        "node-domexception": "1.0.0",
        "web-streams-polyfill": "4.0.0-beta.3"
      },
      "engines": {
        "node": ">= 12.20"
      }
    },
    "node_modules/humanize-ms": {
      "version": "1.2.1",
      "resolved": "https://registry.npmjs.org/humanize-ms/-/humanize-ms-1.2.1.tgz",
      "integrity": "sha512-Fl70vYtsAFb/C06PTS9dZBo7ihau+Tu/DNCk/OyHhea07S+aeMWpFFkUaXRa8fI+ScZbEI8dfSxwY7gxZ9SAVQ==",
      "dependencies": {
        "ms": "^2.0.0"
      }
    },
    "node_modules/is-any-array": {
      "version": "2.0.1",
      "resolved": "https://registry.npmjs.org/is-any-array/-/is-any-array-2.0.1.tgz",
      "integrity": "sha512-UtilS7hLRu++wb/WBAw9bNuP1Eg04Ivn1vERJck8zJthEvXCBEBpGR/33u/xLKWEQf95803oalHrVDptcAvFdQ=="
    },
    "node_modules/is-buffer": {
      "version": "1.1.6",
      "resolved": "https://registry.npmjs.org/is-buffer/-/is-buffer-1.1.6.tgz",
      "integrity": "sha512-NcdALwpXkTm5Zvvbk7owOUSvVvBKDgKP5/ewfXEznmQFfs4ZRmanOeKBTjRVjka3QFoN6XJ+9F3USqfHqTaU5w=="
    },
    "node_modules/js-tiktoken": {
      "version": "1.0.7",
      "resolved": "https://registry.npmjs.org/js-tiktoken/-/js-tiktoken-1.0.7.tgz",
      "integrity": "sha512-biba8u/clw7iesNEWLOLwrNGoBP2lA+hTaBLs/D45pJdUPFXyxD6nhcDVtADChghv4GgyAiMKYMiRx7x6h7Biw==",
      "dependencies": {
        "base64-js": "^1.5.1"
      }
    },
    "node_modules/js-yaml": {
      "version": "4.1.0",
      "resolved": "https://registry.npmjs.org/js-yaml/-/js-yaml-4.1.0.tgz",
      "integrity": "sha512-wpxZs9NoxZaJESJGIZTyDEaYpl0FKSA+FB9aJiyemKhMwkxQg63h4T1KJgUGHpTqPDNRcmmYLugrRjJlBtWvRA==",
      "dependencies": {
        "argparse": "^2.0.1"
      },
      "bin": {
        "js-yaml": "bin/js-yaml.js"
      }
    },
    "node_modules/jsonpointer": {
      "version": "5.0.1",
      "resolved": "https://registry.npmjs.org/jsonpointer/-/jsonpointer-5.0.1.tgz",
      "integrity": "sha512-p/nXbhSEcu3pZRdkW1OfJhpsVtW1gd4Wa1fnQc9YLiTfAjn0312eMKimbdIQzuZl9aa9xUGaRlP9T/CJE/ditQ==",
      "engines": {
        "node": ">=0.10.0"
      }
    },
    "node_modules/langchain": {
      "version": "0.0.165",
      "resolved": "https://registry.npmjs.org/langchain/-/langchain-0.0.165.tgz",
      "integrity": "sha512-CpbNpjwaE+9lzjdw+pZz0VgnRrFivEgr7CVp9dDaAb5JpaJAA4V2v6uQ9ZPN+TSqupTQ79HFn2sfyZVEl2EG7Q==",
      "dependencies": {
        "@anthropic-ai/sdk": "^0.6.2",
        "ansi-styles": "^5.0.0",
        "binary-extensions": "^2.2.0",
        "camelcase": "6",
        "decamelize": "^1.2.0",
        "expr-eval": "^2.0.2",
        "flat": "^5.0.2",
        "js-tiktoken": "^1.0.7",
        "js-yaml": "^4.1.0",
        "jsonpointer": "^5.0.1",
        "langchainhub": "~0.0.6",
        "langsmith": "~0.0.31",
        "ml-distance": "^4.0.0",
        "object-hash": "^3.0.0",
        "openai": "~4.4.0",
        "openapi-types": "^12.1.3",
        "p-queue": "^6.6.2",
        "p-retry": "4",
        "uuid": "^9.0.0",
        "yaml": "^2.2.1",
        "zod": "^3.22.3",
        "zod-to-json-schema": "^3.20.4"
      },
      "engines": {
        "node": ">=18"
      },
      "peerDependencies": {
        "@aws-crypto/sha256-js": "^5.0.0",
        "@aws-sdk/client-bedrock-runtime": "^3.422.0",
        "@aws-sdk/client-dynamodb": "^3.310.0",
        "@aws-sdk/client-kendra": "^3.352.0",
        "@aws-sdk/client-lambda": "^3.310.0",
        "@aws-sdk/client-s3": "^3.310.0",
        "@aws-sdk/client-sagemaker-runtime": "^3.310.0",
        "@aws-sdk/client-sfn": "^3.310.0",
        "@aws-sdk/credential-provider-node": "^3.388.0",
        "@azure/storage-blob": "^12.15.0",
        "@clickhouse/client": "^0.0.14",
        "@cloudflare/ai": "^1.0.12",
        "@elastic/elasticsearch": "^8.4.0",
        "@getmetal/metal-sdk": "*",
        "@getzep/zep-js": "^0.7.0",
        "@gomomento/sdk": "^1.23.0",
        "@google-ai/generativelanguage": "^0.2.1",
        "@google-cloud/storage": "^6.10.1",
        "@huggingface/inference": "^1.5.1",
        "@mozilla/readability": "*",
        "@notionhq/client": "^2.2.10",
        "@opensearch-project/opensearch": "*",
        "@pinecone-database/pinecone": "^1.1.0",
        "@planetscale/database": "^1.8.0",
        "@qdrant/js-client-rest": "^1.2.0",
        "@raycast/api": "^1.55.2",
        "@smithy/eventstream-codec": "^2.0.5",
        "@smithy/protocol-http": "^3.0.6",
        "@smithy/signature-v4": "^2.0.10",
        "@smithy/util-utf8": "^2.0.0",
        "@supabase/postgrest-js": "^1.1.1",
        "@supabase/supabase-js": "^2.10.0",
        "@tensorflow-models/universal-sentence-encoder": "*",
        "@tensorflow/tfjs-converter": "*",
        "@tensorflow/tfjs-core": "*",
        "@upstash/redis": "^1.20.6",
        "@vercel/postgres": "^0.5.0",
        "@writerai/writer-sdk": "^0.40.2",
        "@xata.io/client": "^0.25.1",
        "@xenova/transformers": "^2.5.4",
        "@zilliz/milvus2-sdk-node": ">=2.2.7",
        "apify-client": "^2.7.1",
        "axios": "*",
        "cassandra-driver": "^4.6.4",
        "cheerio": "^1.0.0-rc.12",
        "chromadb": "*",
        "cohere-ai": ">=6.0.0",
        "d3-dsv": "^2.0.0",
        "epub2": "^3.0.1",
        "faiss-node": "^0.3.0",
        "fast-xml-parser": "^4.2.7",
        "firebase-admin": "^11.9.0",
        "google-auth-library": "^8.9.0",
        "googleapis": "^126.0.1",
        "hnswlib-node": "^1.4.2",
        "html-to-text": "^9.0.5",
        "ignore": "^5.2.0",
        "ioredis": "^5.3.2",
        "jsdom": "*",
        "llmonitor": "*",
        "lodash": "^4.17.21",
        "mammoth": "*",
        "mongodb": "^5.2.0",
        "mysql2": "^3.3.3",
        "neo4j-driver": "*",
        "node-llama-cpp": "*",
        "notion-to-md": "^3.1.0",
        "pdf-parse": "1.1.1",
        "peggy": "^3.0.2",
        "pg": "^8.11.0",
        "pg-copy-streams": "^6.0.5",
        "pickleparser": "^0.1.0",
        "playwright": "^1.32.1",
        "portkey-ai": "^0.1.11",
        "puppeteer": "^19.7.2",
        "redis": "^4.6.4",
        "replicate": "^0.18.0",
        "sonix-speech-recognition": "^2.1.1",
        "srt-parser-2": "^1.2.2",
        "typeorm": "^0.3.12",
        "typesense": "^1.5.3",
        "usearch": "^1.1.1",
        "vectordb": "^0.1.4",
        "voy-search": "0.6.2",
        "weaviate-ts-client": "^1.4.0",
        "web-auth-library": "^1.0.3",
        "youtube-transcript": "^1.0.6",
        "youtubei.js": "^5.8.0"
      },
      "peerDependenciesMeta": {
        "@aws-crypto/sha256-js": {
          "optional": true
        },
        "@aws-sdk/client-bedrock-runtime": {
          "optional": true
        },
        "@aws-sdk/client-dynamodb": {
          "optional": true
        },
        "@aws-sdk/client-kendra": {
          "optional": true
        },
        "@aws-sdk/client-lambda": {
          "optional": true
        },
        "@aws-sdk/client-s3": {
          "optional": true
        },
        "@aws-sdk/client-sagemaker-runtime": {
          "optional": true
        },
        "@aws-sdk/client-sfn": {
          "optional": true
        },
        "@aws-sdk/credential-provider-node": {
          "optional": true
        },
        "@azure/storage-blob": {
          "optional": true
        },
        "@clickhouse/client": {
          "optional": true
        },
        "@cloudflare/ai": {
          "optional": true
        },
        "@elastic/elasticsearch": {
          "optional": true
        },
        "@getmetal/metal-sdk": {
          "optional": true
        },
        "@getzep/zep-js": {
          "optional": true
        },
        "@gomomento/sdk": {
          "optional": true
        },
        "@google-ai/generativelanguage": {
          "optional": true
        },
        "@google-cloud/storage": {
          "optional": true
        },
        "@huggingface/inference": {
          "optional": true
        },
        "@mozilla/readability": {
          "optional": true
        },
        "@notionhq/client": {
          "optional": true
        },
        "@opensearch-project/opensearch": {
          "optional": true
        },
        "@pinecone-database/pinecone": {
          "optional": true
        },
        "@planetscale/database": {
          "optional": true
        },
        "@qdrant/js-client-rest": {
          "optional": true
        },
        "@raycast/api": {
          "optional": true
        },
        "@smithy/eventstream-codec": {
          "optional": true
        },
        "@smithy/protocol-http": {
          "optional": true
        },
        "@smithy/signature-v4": {
          "optional": true
        },
        "@smithy/util-utf8": {
          "optional": true
        },
        "@supabase/postgrest-js": {
          "optional": true
        },
        "@supabase/supabase-js": {
          "optional": true
        },
        "@tensorflow-models/universal-sentence-encoder": {
          "optional": true
        },
        "@tensorflow/tfjs-converter": {
          "optional": true
        },
        "@tensorflow/tfjs-core": {
          "optional": true
        },
        "@upstash/redis": {
          "optional": true
        },
        "@vercel/postgres": {
          "optional": true
        },
        "@writerai/writer-sdk": {
          "optional": true
        },
        "@xata.io/client": {
          "optional": true
        },
        "@xenova/transformers": {
          "optional": true
        },
        "@zilliz/milvus2-sdk-node": {
          "optional": true
        },
        "apify-client": {
          "optional": true
        },
        "axios": {
          "optional": true
        },
        "cassandra-driver": {
          "optional": true
        },
        "cheerio": {
          "optional": true
        },
        "chromadb": {
          "optional": true
        },
        "cohere-ai": {
          "optional": true
        },
        "d3-dsv": {
          "optional": true
        },
        "epub2": {
          "optional": true
        },
        "faiss-node": {
          "optional": true
        },
        "fast-xml-parser": {
          "optional": true
        },
        "firebase-admin": {
          "optional": true
        },
        "google-auth-library": {
          "optional": true
        },
        "googleapis": {
          "optional": true
        },
        "hnswlib-node": {
          "optional": true
        },
        "html-to-text": {
          "optional": true
        },
        "ignore": {
          "optional": true
        },
        "ioredis": {
          "optional": true
        },
        "jsdom": {
          "optional": true
        },
        "llmonitor": {
          "optional": true
        },
        "lodash": {
          "optional": true
        },
        "mammoth": {
          "optional": true
        },
        "mongodb": {
          "optional": true
        },
        "mysql2": {
          "optional": true
        },
        "neo4j-driver": {
          "optional": true
        },
        "node-llama-cpp": {
          "optional": true
        },
        "notion-to-md": {
          "optional": true
        },
        "pdf-parse": {
          "optional": true
        },
        "peggy": {
          "optional": true
        },
        "pg": {
          "optional": true
        },
        "pg-copy-streams": {
          "optional": true
        },
        "pickleparser": {
          "optional": true
        },
        "playwright": {
          "optional": true
        },
        "portkey-ai": {
          "optional": true
        },
        "puppeteer": {
          "optional": true
        },
        "redis": {
          "optional": true
        },
        "replicate": {
          "optional": true
        },
        "sonix-speech-recognition": {
          "optional": true
        },
        "srt-parser-2": {
          "optional": true
        },
        "typeorm": {
          "optional": true
        },
        "typesense": {
          "optional": true
        },
        "usearch": {
          "optional": true
        },
        "vectordb": {
          "optional": true
        },
        "voy-search": {
          "optional": true
        },
        "weaviate-ts-client": {
          "optional": true
        },
        "web-auth-library": {
          "optional": true
        },
        "youtube-transcript": {
          "optional": true
        },
        "youtubei.js": {
          "optional": true
        }
      }
    },
    "node_modules/langchainhub": {
      "version": "0.0.6",
      "resolved": "https://registry.npmjs.org/langchainhub/-/langchainhub-0.0.6.tgz",
      "integrity": "sha512-SW6105T+YP1cTe0yMf//7kyshCgvCTyFBMTgH2H3s9rTAR4e+78DA/BBrUL/Mt4Q5eMWui7iGuAYb3pgGsdQ9w=="
    },
    "node_modules/langsmith": {
      "version": "0.0.42",
      "resolved": "https://registry.npmjs.org/langsmith/-/langsmith-0.0.42.tgz",
      "integrity": "sha512-sFuN+e7E+pPBIRaRgFqZh/BRBWNHTZNAwi6uj4kydQawooCZYoJmM5snOkiQrhVSvAhgu6xFhLvmfvkPcKzD7w==",
      "dependencies": {
        "@types/uuid": "^9.0.1",
        "commander": "^10.0.1",
        "p-queue": "^6.6.2",
        "p-retry": "4",
        "uuid": "^9.0.0"
      },
      "bin": {
        "langsmith": "dist/cli/main.cjs"
      }
    },
    "node_modules/md5": {
      "version": "2.3.0",
      "resolved": "https://registry.npmjs.org/md5/-/md5-2.3.0.tgz",
      "integrity": "sha512-T1GITYmFaKuO91vxyoQMFETst+O71VUPEU3ze5GNzDm0OWdP8v1ziTaAEPUr/3kLsY3Sftgz242A1SetQiDL7g==",
      "dependencies": {
        "charenc": "0.0.2",
        "crypt": "0.0.2",
        "is-buffer": "~1.1.6"
      }
    },
    "node_modules/mime-db": {
      "version": "1.52.0",
      "resolved": "https://registry.npmjs.org/mime-db/-/mime-db-1.52.0.tgz",
      "integrity": "sha512-sPU4uV7dYlvtWJxwwxHD0PuihVNiE7TyAbQ5SWxDCB9mUYvOgroQOwYQQOKPJ8CIbE+1ETVlOoK1UC2nU3gYvg==",
      "engines": {
        "node": ">= 0.6"
      }
    },
    "node_modules/mime-types": {
      "version": "2.1.35",
      "resolved": "https://registry.npmjs.org/mime-types/-/mime-types-2.1.35.tgz",
      "integrity": "sha512-ZDY+bPm5zTTF+YpCrAU9nK0UgICYPT0QtT1NZWFv4s++TNkcgVaT0g6+4R2uI4MjQjzysHB1zxuWL50hzaeXiw==",
      "dependencies": {
        "mime-db": "1.52.0"
      },
      "engines": {
        "node": ">= 0.6"
      }
    },
    "node_modules/ml-array-mean": {
      "version": "1.1.6",
      "resolved": "https://registry.npmjs.org/ml-array-mean/-/ml-array-mean-1.1.6.tgz",
      "integrity": "sha512-MIdf7Zc8HznwIisyiJGRH9tRigg3Yf4FldW8DxKxpCCv/g5CafTw0RRu51nojVEOXuCQC7DRVVu5c7XXO/5joQ==",
      "dependencies": {
        "ml-array-sum": "^1.1.6"
      }
    },
    "node_modules/ml-array-sum": {
      "version": "1.1.6",
      "resolved": "https://registry.npmjs.org/ml-array-sum/-/ml-array-sum-1.1.6.tgz",
      "integrity": "sha512-29mAh2GwH7ZmiRnup4UyibQZB9+ZLyMShvt4cH4eTK+cL2oEMIZFnSyB3SS8MlsTh6q/w/yh48KmqLxmovN4Dw==",
      "dependencies": {
        "is-any-array": "^2.0.0"
      }
    },
    "node_modules/ml-distance": {
      "version": "4.0.1",
      "resolved": "https://registry.npmjs.org/ml-distance/-/ml-distance-4.0.1.tgz",
      "integrity": "sha512-feZ5ziXs01zhyFUUUeZV5hwc0f5JW0Sh0ckU1koZe/wdVkJdGxcP06KNQuF0WBTj8FttQUzcvQcpcrOp/XrlEw==",
      "dependencies": {
        "ml-array-mean": "^1.1.6",
        "ml-distance-euclidean": "^2.0.0",
        "ml-tree-similarity": "^1.0.0"
      }
    },
    "node_modules/ml-distance-euclidean": {
      "version": "2.0.0",
      "resolved": "https://registry.npmjs.org/ml-distance-euclidean/-/ml-distance-euclidean-2.0.0.tgz",
      "integrity": "sha512-yC9/2o8QF0A3m/0IXqCTXCzz2pNEzvmcE/9HFKOZGnTjatvBbsn4lWYJkxENkA4Ug2fnYl7PXQxnPi21sgMy/Q=="
    },
    "node_modules/ml-tree-similarity": {
      "version": "1.0.0",
      "resolved": "https://registry.npmjs.org/ml-tree-similarity/-/ml-tree-similarity-1.0.0.tgz",
      "integrity": "sha512-XJUyYqjSuUQkNQHMscr6tcjldsOoAekxADTplt40QKfwW6nd++1wHWV9AArl0Zvw/TIHgNaZZNvr8QGvE8wLRg==",
      "dependencies": {
        "binary-search": "^1.3.5",
        "num-sort": "^2.0.0"
      }
    },
    "node_modules/ms": {
      "version": "2.1.3",
      "resolved": "https://registry.npmjs.org/ms/-/ms-2.1.3.tgz",
      "integrity": "sha512-6FlzubTLZG3J2a/NVCAleEhjzq5oxgHyaCU9yYXvcLsvoVaHJq/s5xXI6/XXP6tz7R9xAOtHnSO/tXtF3WRTlA=="
    },
    "node_modules/node-domexception": {
      "version": "1.0.0",
      "resolved": "https://registry.npmjs.org/node-domexception/-/node-domexception-1.0.0.tgz",
      "integrity": "sha512-/jKZoMpw0F8GRwl4/eLROPA3cfcXtLApP0QzLmUT/HuPCZWyB7IY9ZrMeKw2O/nFIqPQB3PVM9aYm0F312AXDQ==",
      "funding": [
        {
          "type": "github",
          "url": "https://github.com/sponsors/jimmywarting"
        },
        {
          "type": "github",
          "url": "https://paypal.me/jimmywarting"
        }
      ],
      "engines": {
        "node": ">=10.5.0"
      }
    },
    "node_modules/node-fetch": {
      "version": "2.7.0",
      "resolved": "https://registry.npmjs.org/node-fetch/-/node-fetch-2.7.0.tgz",
      "integrity": "sha512-c4FRfUm/dbcWZ7U+1Wq0AwCyFL+3nt2bEw05wfxSz+DWpWsitgmSgYmy2dQdWyKC1694ELPqMs/YzUSNozLt8A==",
      "dependencies": {
        "whatwg-url": "^5.0.0"
      },
      "engines": {
        "node": "4.x || >=6.0.0"
      },
      "peerDependencies": {
        "encoding": "^0.1.0"
      },
      "peerDependenciesMeta": {
        "encoding": {
          "optional": true
        }
      }
    },
    "node_modules/num-sort": {
      "version": "2.1.0",
      "resolved": "https://registry.npmjs.org/num-sort/-/num-sort-2.1.0.tgz",
      "integrity": "sha512-1MQz1Ed8z2yckoBeSfkQHHO9K1yDRxxtotKSJ9yvcTUUxSvfvzEq5GwBrjjHEpMlq/k5gvXdmJ1SbYxWtpNoVg==",
      "engines": {
        "node": ">=8"
      },
      "funding": {
        "url": "https://github.com/sponsors/sindresorhus"
      }
    },
    "node_modules/object-hash": {
      "version": "3.0.0",
      "resolved": "https://registry.npmjs.org/object-hash/-/object-hash-3.0.0.tgz",
      "integrity": "sha512-RSn9F68PjH9HqtltsSnqYC1XXoWe9Bju5+213R98cNGttag9q9yAOTzdbsqvIa7aNm5WffBZFpWYr2aWrklWAw==",
      "engines": {
        "node": ">= 6"
      }
    },
    "node_modules/openai": {
      "version": "4.4.0",
      "resolved": "https://registry.npmjs.org/openai/-/openai-4.4.0.tgz",
      "integrity": "sha512-JN0t628Kh95T0IrXl0HdBqnlJg+4Vq0Bnh55tio+dfCnyzHvMLiWyCM9m726MAJD2YkDU4/8RQB6rNbEq9ct2w==",
      "dependencies": {
        "@types/node": "^18.11.18",
        "@types/node-fetch": "^2.6.4",
        "abort-controller": "^3.0.0",
        "agentkeepalive": "^4.2.1",
        "digest-fetch": "^1.3.0",
        "form-data-encoder": "1.7.2",
        "formdata-node": "^4.3.2",
        "node-fetch": "^2.6.7"
      },
      "bin": {
        "openai": "bin/cli"
      }
    },
    "node_modules/openapi-types": {
      "version": "12.1.3",
      "resolved": "https://registry.npmjs.org/openapi-types/-/openapi-types-12.1.3.tgz",
      "integrity": "sha512-N4YtSYJqghVu4iek2ZUvcN/0aqH1kRDuNqzcycDxhOUpg7GdvLa2F3DgS6yBNhInhv2r/6I0Flkn7CqL8+nIcw=="
    },
    "node_modules/p-finally": {
      "version": "1.0.0",
      "resolved": "https://registry.npmjs.org/p-finally/-/p-finally-1.0.0.tgz",
      "integrity": "sha512-LICb2p9CB7FS+0eR1oqWnHhp0FljGLZCWBE9aix0Uye9W8LTQPwMTYVGWQWIw9RdQiDg4+epXQODwIYJtSJaow==",
      "engines": {
        "node": ">=4"
      }
    },
    "node_modules/p-queue": {
      "version": "6.6.2",
      "resolved": "https://registry.npmjs.org/p-queue/-/p-queue-6.6.2.tgz",
      "integrity": "sha512-RwFpb72c/BhQLEXIZ5K2e+AhgNVmIejGlTgiB9MzZ0e93GRvqZ7uSi0dvRF7/XIXDeNkra2fNHBxTyPDGySpjQ==",
      "dependencies": {
        "eventemitter3": "^4.0.4",
        "p-timeout": "^3.2.0"
      },
      "engines": {
        "node": ">=8"
      },
      "funding": {
        "url": "https://github.com/sponsors/sindresorhus"
      }
    },
    "node_modules/p-retry": {
      "version": "4.6.2",
      "resolved": "https://registry.npmjs.org/p-retry/-/p-retry-4.6.2.tgz",
      "integrity": "sha512-312Id396EbJdvRONlngUx0NydfrIQ5lsYu0znKVUzVvArzEIt08V1qhtyESbGVd1FGX7UKtiFp5uwKZdM8wIuQ==",
      "dependencies": {
        "@types/retry": "0.12.0",
        "retry": "^0.13.1"
      },
      "engines": {
        "node": ">=8"
      }
    },
    "node_modules/p-timeout": {
      "version": "3.2.0",
      "resolved": "https://registry.npmjs.org/p-timeout/-/p-timeout-3.2.0.tgz",
      "integrity": "sha512-rhIwUycgwwKcP9yTOOFK/AKsAopjjCakVqLHePO3CC6Mir1Z99xT+R63jZxAT5lFZLa2inS5h+ZS2GvR99/FBg==",
      "dependencies": {
        "p-finally": "^1.0.0"
      },
      "engines": {
        "node": ">=8"
      }
    },
    "node_modules/retry": {
      "version": "0.13.1",
      "resolved": "https://registry.npmjs.org/retry/-/retry-0.13.1.tgz",
      "integrity": "sha512-XQBQ3I8W1Cge0Seh+6gjj03LbmRFWuoszgK9ooCpwYIrhhoO80pfq4cUkU5DkknwfOfFteRwlZ56PYOGYyFWdg==",
      "engines": {
        "node": ">= 4"
      }
    },
    "node_modules/tr46": {
      "version": "0.0.3",
      "resolved": "https://registry.npmjs.org/tr46/-/tr46-0.0.3.tgz",
      "integrity": "sha512-N3WMsuqV66lT30CrXNbEjx4GEwlow3v6rr4mCcv6prnfwhS01rkgyFdjPNBYd9br7LpXV1+Emh01fHnq2Gdgrw=="
    },
    "node_modules/typescript": {
      "version": "5.2.2",
      "resolved": "https://registry.npmjs.org/typescript/-/typescript-5.2.2.tgz",
      "integrity": "sha512-mI4WrpHsbCIcwT9cF4FZvr80QUeKvsUsUvKDoR+X/7XHQH98xYD8YHZg7ANtz2GtZt/CBq2QJ0thkGJMHfqc1w==",
      "dev": true,
      "bin": {
        "tsc": "bin/tsc",
        "tsserver": "bin/tsserver"
      },
      "engines": {
        "node": ">=14.17"
      }
    },
    "node_modules/uuid": {
      "version": "9.0.1",
      "resolved": "https://registry.npmjs.org/uuid/-/uuid-9.0.1.tgz",
      "integrity": "sha512-b+1eJOlsR9K8HJpow9Ok3fiWOWSIcIzXodvv0rQjVoOVNpWMpxf1wZNpt4y9h10odCNrqnYp1OBzRktckBe3sA==",
      "funding": [
        "https://github.com/sponsors/broofa",
        "https://github.com/sponsors/ctavan"
      ],
      "bin": {
        "uuid": "dist/bin/uuid"
      }
    },
    "node_modules/web-streams-polyfill": {
      "version": "4.0.0-beta.3",
      "resolved": "https://registry.npmjs.org/web-streams-polyfill/-/web-streams-polyfill-4.0.0-beta.3.tgz",
      "integrity": "sha512-QW95TCTaHmsYfHDybGMwO5IJIM93I/6vTRk+daHTWFPhwh+C8Cg7j7XyKrwrj8Ib6vYXe0ocYNrmzY4xAAN6ug==",
      "engines": {
        "node": ">= 14"
      }
    },
    "node_modules/webidl-conversions": {
      "version": "3.0.1",
      "resolved": "https://registry.npmjs.org/webidl-conversions/-/webidl-conversions-3.0.1.tgz",
      "integrity": "sha512-2JAn3z8AR6rjK8Sm8orRC0h/bcl/DqL7tRPdGZ4I1CjdF+EaMLmYxBHyXuKL849eucPFhvBoxMsflfOb8kxaeQ=="
    },
    "node_modules/whatwg-url": {
      "version": "5.0.0",
      "resolved": "https://registry.npmjs.org/whatwg-url/-/whatwg-url-5.0.0.tgz",
      "integrity": "sha512-saE57nupxk6v3HY35+jzBwYa0rKSy0XR8JSxZPwgLr7ys0IBzhGviA1/TUGJLmSVqs8pb9AnvICXEuOHLprYTw==",
      "dependencies": {
        "tr46": "~0.0.3",
        "webidl-conversions": "^3.0.0"
      }
    },
    "node_modules/yaml": {
      "version": "2.3.2",
      "resolved": "https://registry.npmjs.org/yaml/-/yaml-2.3.2.tgz",
      "integrity": "sha512-N/lyzTPaJasoDmfV7YTrYCI0G/3ivm/9wdG0aHuheKowWQwGTsK0Eoiw6utmzAnI6pkJa0DUVygvp3spqqEKXg==",
      "engines": {
        "node": ">= 14"
      }
    },
    "node_modules/zod": {
      "version": "3.22.4",
      "resolved": "https://registry.npmjs.org/zod/-/zod-3.22.4.tgz",
      "integrity": "sha512-iC+8Io04lddc+mVqQ9AZ7OQ2MrUKGN+oIQyq1vemgt46jwCwLfhq7/pwnBnNXXXZb8VTVLKwp9EDkx+ryxIWmg==",
      "funding": {
        "url": "https://github.com/sponsors/colinhacks"
      }
    },
    "node_modules/zod-to-json-schema": {
      "version": "3.21.4",
      "resolved": "https://registry.npmjs.org/zod-to-json-schema/-/zod-to-json-schema-3.21.4.tgz",
      "integrity": "sha512-fjUZh4nQ1s6HMccgIeE0VP4QG/YRGPmyjO9sAh890aQKPEk3nqbfUXhMFaC+Dr5KvYBm8BCyvfpZf2jY9aGSsw==",
      "peerDependencies": {
        "zod": "^3.21.4"
      }
    }
  }
 }
--- a/examples/langchain-typescript-simple/package.json
+++ b/examples/langchain-typescript-simple/package.json
@@ -0,0 +1,8 @@
 {
  "devDependencies": {
    "typescript": "^5.2.2"
  },
  "dependencies": {
    "langchain": "^0.0.165"
  }
 }
--- a/examples/midjourney-prompter/Modelfile
+++ b/examples/midjourney-prompter/Modelfile
@@ -1,8 +0,0 @@
 # Modelfile for creating a Midjourney prompts from a topic
 # This prompt was adapted from the original at https://www.greataiprompts.com/guide/midjourney/best-chatgpt-prompt-for-midjourney/
 # Run `ollama create mj -f ./Modelfile` and then `ollama run mj` and enter a topic
 FROM nous-hermes
 SYSTEM """
 Embrace your role as an AI-powered creative assistant, employing Midjourney to manifest compelling AI-generated art. I will outline a specific image concept, and in response, you must produce an exhaustive, multifaceted prompt for Midjourney, ensuring every detail of the original concept is represented in your instructions. Midjourney doesn't do well with text, so after the prompt, give me instructions that I can use to create the titles in a image editor.
 """
--- a/examples/modelfile-10tweets/Modelfile
+++ b/examples/modelfile-10tweets/Modelfile
@@ -0,0 +1,7 @@
 # Modelfile for creating a list of ten tweets from a topic
 # Run `ollama create 10tweets -f ./Modelfile` and then `ollama run 10tweets` and enter a topic
 FROM llama2
 SYSTEM """
 You are a content marketer who needs to come up with 10 short but succinct tweets. The answer should be a list of ten tweets. Each tweet can have a maximum of 280 characters and should include hashtags. Each user input will be a subject and you should expand it in ten creative ways. Never stop after just one tweet. Always include ten. 
 """
--- a/examples/modelfile-10tweets/README.md
+++ b/examples/modelfile-10tweets/README.md
@@ -0,0 +1,23 @@
 # Ten Tweets Modelfile
 This is a simple modelfile that generates ten tweets based off any topic.
 ```bash
 ollama create tentweets
 ollama run tentweets
 >>> underwater basketweaving
 Great! Here are ten creative tweets about underwater basketweaving:
 1. "Just discovered the ultimate stress-reliever: Underwater basketweaving! 🌊🧵 #UnderwaterBasketweaving #StressRelief"
 2. "Who needs meditation when you can do underwater basketweaving? 😴👀 #PeacefulDistraction #UnderwaterBasketweaving"
 3. "Just spent an hour in the pool and still managed to knot my basket. Goal: untangle it before next session. 💪🏽 #ChallengeAccepted #UnderwaterBasketweaving"
 4. "When life gives you lemons, make underwater basketweaving! 🍋🧵 #LemonadeLife #UnderwaterBasketweaving"
 5. "Just realized my underwater basketweaving skills could come in handy during a zombie apocalypse. 😂🧡 #SurvivalTips #UnderwaterBasketweaving"
 6. "I'm not lazy, I'm just conserving energy for my next underwater basketweaving session. 😴💤 #LazyDay #UnderwaterBasketweaving"
 7. "Just found my inner peace while doing underwater basketweaving. It's like meditation, but with knots! 🙏🧵 #Mindfulness #UnderwaterBasketweaving"
 8. "Why study for exams when you can do underwater basketweaving and forget all your worries? 😜🧵 #ProcrastinationStation #UnderwaterBasketweaving"
 9. "Just had to cut my underwater basketweaving session short due to a sudden urge to breathe. 🤯🌊 #AquaticAdventures #UnderwaterBasketweaving"
 10. "I'm not sure what's more impressive: my underwater basketweaving skills or the fact that I didn't drown trying to make this tweet. 😅🧵 #Accomplishment 
 #UnderwaterBasketweaving"
 ```
--- a/examples/modelfile-devopsengineer/Modelfile
+++ b/examples/modelfile-devopsengineer/Modelfile
@@ -0,0 +1,8 @@
 # Modelfile for creating a devops engineer assistant
 # Run `ollama create devops-engineer -f ./Modelfile` and then `ollama run devops-engineer` and enter a topic
 FROM mistral
 PARAMETER temperature 1
 SYSTEM """
 You are a senior devops engineer, acting as an assistant. You offer help with cloud technologies like: Terraform, AWS, kubernetes, python. You answer with code examples when possible
 """
--- a/examples/modelfile-devopsengineer/README.md
+++ b/examples/modelfile-devopsengineer/README.md
@@ -0,0 +1,40 @@
 # DevOps Engineer Modelfile
 This is a simple modelfile that takes will respond to any questions that a senior DevOps engineer should know how to solve.
 ````text
 ollama create devopsengineer
 ollama run devopsengineer
 >>> how can i proxy all connections from inside the office through a single server that caches large files. Provide a config file for nginx to achieve this?
 Here is an example configuration file for Nginx that can be used to proxy all connections from inside the office network to a centralized cache server:
 ```perl
 http {
    upstream cache_server {
        server <cache_server_ip>;
        server <cache_server_ip>;
        server <cache_server_ip>;
    }
    server {
        listen 80;
        server_name localhost;
        location / {
            proxy_pass http://cache_server;
            proxy_set_header Host $host;
            proxy_set_header X-Real-IP $remote_addr;
            proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        }
    }
 }
 ```
 In this configuration, we define an `upstream` block that lists the IP addresses of the cache servers. We then define a server block that listens on port 80 and routes all 
 traffic to the `/` location to the cache servers using the `proxy_pass` directive. The `proxy_set_header` directives are used to preserve the source IP address of the client
 request when forwarding it to the cache server.
 To use this configuration, you would need to replace the placeholder `<cache_server_ip>` with the actual IP addresses of your cache servers. You would also need to make sure
 that the cache servers are configured to accept incoming connections from the Nginx server and handle requests for files.
 ````
--- a/examples/modelfile-mario/Modelfile
+++ b/examples/modelfile-mario/Modelfile
--- a/examples/modelfile-mario/logo.png
+++ b/examples/modelfile-mario/logo.png
--- a/examples/modelfile-mario/readme.md
+++ b/examples/modelfile-mario/readme.md
--- a/examples/modelfile-midjourney/Modelfile
+++ b/examples/modelfile-midjourney/Modelfile
@@ -0,0 +1,11 @@
 # Modelfile for creating a Midjourney prompts from a topic
 # This prompt was adapted from the original at https://www.greataiprompts.com/guide/midjourney/best-chatgpt-prompt-for-midjourney/
 # Run `ollama create mj -f ./Modelfile` and then `ollama run mj` and enter a topic
 FROM zephyr
 PARAMETER temperature 0.8
 PARAMETER top_k 500
 PARAMETER top_p 0.9
 SYSTEM """
 Embrace your role as a creative illustrator. Based on a concept provided, you must produce a single paragraph with a multifaceted description of an image, ensuring significant details of the concept and more is represented in your instructions. You do not need to write complete sentences but rather short concepts with the following information: the level of detail that should be represented, an artistic style and maybe a specific name of a painter or illustrator, the ideal color pallete, lighting, mood, perspective, the setting, time of day, weather, the season, the time period, location, materials, the textures, patterns, lines, brushstrokes, techniques, the medium, the genre, the rendering style. Don't include everything and keep the description length under 250 words. 
 """
--- a/examples/modelfile-midjourney/README.md
+++ b/examples/modelfile-midjourney/README.md
@@ -0,0 +1,11 @@
 # Midjourney Prompt Generator Modelfile
 This simple modelfile will help create a prompt to feed to Midjourney.
 ```text
 ollama create midjourney
 ollama run midjourney
 >>> a sports car in the mountains. 
 A sleek, high-performance automobile cuts through a serpentine mountain landscape. The concept is a classic illustration of speed and power, depicted in the style of pop art by Andy Warhol. The color palette is dominated by bold, primary hues of red, blue, and yellow, with striking accent colors of white, black, and metallic shades. The lighting is bright and focused, casting sharp shadows on the rugged terrain. A sense of excitement and anticipation permeates throughout the scene, as the car navigates a treacherous course through the winding road. The perspective is low, allowing for a full view of the vehicle's sleek lines and intricate details. The setting takes place in the afternoon during a sunny day in autumn, as evidenced by the vibrant foliage on the mountainside. The time period is modern, with nods to classic car design. The materials are primarily digital, allowing for smooth curves and sharp contrasts. The textures are sleek and polished, with meticulously detailed lines and brushstrokes that accentuate the car's aerodynamic design. The patterns consist of geometric shapes and bold stripes, adding to the car's dynamic appeal. The genre is modern realism, with a focus on precision and detail. The rendering style is highly technical, capturing the nuances and subtleties of the vehicle and its surroundings in breathtaking detail.
 ```
--- a/examples/modelfile-recipemaker/Modelfile
+++ b/examples/modelfile-recipemaker/Modelfile
--- a/examples/modelfile-recipemaker/README.md
+++ b/examples/modelfile-recipemaker/README.md
@@ -0,0 +1,20 @@
 # Recipe Maker Modelfile 
 Simple modelfile to generate a recipe from a short list of ingredients.
 ```
 ollama create recipemaker
 ollama run recipemaker
 >>> chilli pepper, white chocolate, kale
 Ingredients:
 - 1 small chili pepper
 - 4 squares of white chocolate
 - handful of kale leaves
 Instructions:
 1. In a blender or food processor, puree the chilies and white chocolate until smooth.
 2. Add the chopped kale leaves to the blender and pulse until well combined.
 3. Serve immediately as a dip for crackers or use it as an ingredient in your favorite recipe. The mixture of spicy chili pepper with sweet white chocolate and nutritious 
 kale will make your taste buds dance with delight!
 ```
--- a/examples/modelfile-sentiments/Modelfile
+++ b/examples/modelfile-sentiments/Modelfile
@@ -0,0 +1,28 @@
 # Modelfile for creating a sentiment analyzer. 
 # Run `ollama create sentiments -f pathtofile` and then `ollama run sentiments` and enter a topic
 FROM orca
 TEMPLATE """
 {{- if .First }}
 ### System:
 {{ .System }}
 {{- end }}
 ### User: 
 I hate it when my phone dies
 ### Response: 
 NEGATIVE
 ### User: 
 He is awesome
 ### Response: 
 POSITIVE
 ### User: 
 This is the link to the article
 ### Response: 
 NEUTRAL
 ### User:
 {{ .Prompt }}
 ### Response:
 """
 SYSTEM """You are a sentiment analyzer. You will receive text and output only one word, either POSITIVE or NEGATIVE or NEUTRAL, depending on the sentiment of the text."""
--- a/examples/modelfile-sentiments/Readme.md
+++ b/examples/modelfile-sentiments/Readme.md
@@ -0,0 +1,25 @@
 # Sentiments Modelfile
 This is a simple sentiments analyzer using the Orca model. When you pull Orca from the registry, it has a Template already defined that looks like this:
 ```Modelfile
 {{- if .First }}
 ### System:
 {{ .System }}
 {{- end }}
 ### User:
 {{ .Prompt }}
 ### Response:
 ```
 If we just wanted to have the text:
 ```Plaintext
 You are a sentiment analyzer. You will receive text and output only one word, either POSITIVE or NEGATIVE or NEUTRAL, depending on the sentiment of the text.
 ```
 then we could have put this in a SYSTEM block. But we want to provide examples which require updating the full Template. Any Modelfile you create will inherit all the settings from the source model. But in this example, we are overriding the Template.
 When providing examples for the input and output, you should include the way the model usually provides information. Since the Orca model expects a user prompt to appear after ### User: and the response is after ### Response, we should format our examples like that as well. If we were using the Llama 2 model, the format would be a bit different.
--- a/examples/modelfile-tweetwriter/Modelfile
+++ b/examples/modelfile-tweetwriter/Modelfile
@@ -3,5 +3,5 @@
 FROM nous-hermes
 SYSTEM """
-You are a content marketer who needs to come up with a short but succinct tweet. Make sure to include the appropriate hashtags and links. Sometimes when appropriate, describe a meme that can be includes as well. All answers should be in the form of a tweet which has a max size of 280 characters. Every instruction will be the topic to create a tweet about.
+You are a content marketer who needs to come up with a short but succinct tweet. Make sure to include the appropriate hashtags and links. Sometimes when appropriate, describe a meme that can be included as well. All answers should be in the form of a tweet which has a max size of 280 characters. Every instruction will be the topic to create a tweet about.
 """
--- a/examples/python-dockerit/Modelfile
+++ b/examples/python-dockerit/Modelfile
@@ -0,0 +1,20 @@
 FROM mistral
 SYSTEM """
 You are an experienced Devops engineer focused on docker. When given specifications for a particular need or application you know the best way to host that within a docker container. For instance if someone tells you they want an nginx server to host files located at /web you will answer as follows
 ---start
 FROM nginx:alpine
 COPY /myweb /usr/share/nginx/html
 EXPOSE 80
 ---end
 Notice that the answer you should give is just the contents of the dockerfile with no explanation and there are three dashes and the word start at the beginning and 3 dashes and the word end. The full output can be piped into a file and run as is. Here is another example. The user will ask to launch a Postgres server with a password of abc123. And the response should be
 ---start
 FROM postgres:latest
 ENV POSTGRES_PASSWORD=abc123
 EXPOSE 5432
 ---end
 Again it's just the contents of the dockerfile and nothing else.
 """
--- a/examples/python-dockerit/README.md
+++ b/examples/python-dockerit/README.md
@@ -0,0 +1,15 @@
 # DockerIt
 DockerIt is a tool to help you build and run your application in a Docker container. It consists of a model that defines the system prompt and model weights to use, along with a python script to then build the container and run the image automatically. 
 ## Caveats
 This is an simple example. It's assuming the Dockerfile content generated is going to work. In many cases, even with simple web servers, it fails when trying to copy files that don't exist. It's simply an example of what you could possibly do.
 ## Example Usage
 ```bash
 > python3 ./dockerit.py "simple postgres server with admin password set to 123"
 Enter the name of the image: matttest
 Container named happy_keller  started with id:  7c201bb6c30f02b356ddbc8e2a5af9d7d7d7b8c228519c9a501d15c0bd9d6b3e
 ```
--- a/examples/python-dockerit/dockerit.py
+++ b/examples/python-dockerit/dockerit.py
@@ -0,0 +1,17 @@
 import requests, json, docker, io, sys
 inputDescription = " ".join(sys.argv[1:])
 imageName = input("Enter the name of the image: ")
 client = docker.from_env()
 s = requests.Session()
 output=""
 with s.post('http://localhost:11434/api/generate', json={'model': 'dockerit', 'prompt': inputDescription}, stream=True) as r:
  for line in r.iter_lines():
    if line:
      j = json.loads(line)
      if "response" in j:
        output = output +j["response"]
 output = output[output.find("---start")+9:output.find("---end")-1]
 f = io.BytesIO(bytes(output, 'utf-8'))
 client.images.build(fileobj=f, tag=imageName)
 container = client.containers.run(imageName, detach=True)
 print("Container named", container.name, " started with id: ",container.id)
--- a/examples/python-dockerit/requirements.txt
+++ b/examples/python-dockerit/requirements.txt
@@ -0,0 +1 @@
 docker
--- a/examples/python-simplegenerate/client.py
+++ b/examples/python-simplegenerate/client.py
@@ -0,0 +1,38 @@
 import json
 import requests
 # NOTE: ollama must be running for this to work, start the ollama app or run `ollama serve`
 model = 'llama2' # TODO: update this for whatever model you wish to use
 def generate(prompt, context):
    r = requests.post('http://localhost:11434/api/generate',
                      json={
                          'model': model,
                          'prompt': prompt,
                          'context': context,
                      },
                      stream=True)
    r.raise_for_status()
    for line in r.iter_lines():
        body = json.loads(line)
        response_part = body.get('response', '')
        # the response streams one token at a time, print that as we recieve it
        print(response_part, end='', flush=True)
        if 'error' in body:
            raise Exception(body['error'])
        if body.get('done', False):
            return body['context']
 def main():
    context = [] # the context stores a conversation history, you can use this to make the model more context aware
    while True:
        user_input = input("Enter a prompt: ")
        print()
        context = generate(user_input, context)
        print()
 if __name__ == "__main__":
    main()
--- a/examples/typescript-mentors/.gitignore
+++ b/examples/typescript-mentors/.gitignore
@@ -0,0 +1,2 @@
 node_modules
 package-lock.json
--- a/examples/typescript-mentors/README.md
+++ b/examples/typescript-mentors/README.md
@@ -0,0 +1,21 @@
 # Ask the Mentors
 This example demonstrates how one would create a set of 'mentors' you can have a conversation with. The mentors are generated using the `character-generator.ts` file. This will use **Stable Beluga 70b** to create a bio and list of verbal ticks and common phrases used by each person. Then `mentors.ts` will take a question, and choose three of the 'mentors' and start a conversation with them. Occasionally, they will talk to each other, and other times they will just deliver a set of monologues. It's fun to see what they do and say.
 ## Usage
 ```bash
 ts-node ./character-generator.ts "Lorne Greene"
 ```
 This will create `lornegreene/Modelfile`. Now you can create a model with this command:
 ```bash
 ollama create lornegreene -f lornegreene/Modelfile
 ```
 If you want to add your own mentors, you will have to update the code to look at your namespace instead of **mattw**. Also set the list of mentors to include yours.
 ```bash
 ts-node ./mentors.ts "What is a Jackalope?"
 ```
--- a/examples/typescript-mentors/character-generator.ts
+++ b/examples/typescript-mentors/character-generator.ts
@@ -0,0 +1,26 @@
 import { Ollama } from 'ollama-node'
 import fs from 'fs';
 import path from 'path';
 async function characterGenerator() {
  const character = process.argv[2];
  console.log(`You are creating a character for ${character}.`);
  const foldername = character.replace(/\s/g, '').toLowerCase();
  const directory = path.join(__dirname, foldername);
  if (!fs.existsSync(directory)) {
    fs.mkdirSync(directory, { recursive: true });
  }
  const ollama = new Ollama();
  ollama.setModel("stablebeluga2:70b-q4_K_M");
  const bio = await ollama.generate(`create a bio of ${character} in a single long paragraph. Instead of saying '${character} is...' or '${character} was...' use language like 'You are...' or 'You were...'. Then create a paragraph describing the speaking mannerisms and style of ${character}. Don't include anything about how ${character} looked or what they sounded like, just focus on the words they said. Instead of saying '${character} would say...' use language like 'You should say...'. If you use quotes, always use single quotes instead of double quotes. If there are any specific words or phrases you used a lot, show how you used them. `);
  const thecontents = `FROM llama2\nSYSTEM """\n${bio.response.replace(/(\r\n|\n|\r)/gm, " ").replace('would', 'should')} All answers to questions should be related back to what you are most known for.\n"""`;
  fs.writeFile(path.join(directory, 'Modelfile'), thecontents, (err: any) => {
    if (err) throw err;
    console.log('The file has been saved!');
  });
 }
 characterGenerator();
--- a/examples/typescript-mentors/mentors.ts
+++ b/examples/typescript-mentors/mentors.ts
@@ -0,0 +1,59 @@
 import { Ollama } from 'ollama-node';
 const mentorCount = 3;
 const ollama = new Ollama();
 function getMentors(): string[] {
  const mentors = ['Gary Vaynerchuk', 'Kanye West', 'Martha Stewart', 'Neil deGrasse Tyson', 'Owen Wilson', 'Ronald Reagan', 'Donald Trump', 'Barack Obama', 'Jeff Bezos'];
  const chosenMentors: string[] = [];
  for (let i = 0; i < mentorCount; i++) {
    const mentor = mentors[Math.floor(Math.random() * mentors.length)];
    chosenMentors.push(mentor);
    mentors.splice(mentors.indexOf(mentor), 1);
  }
  return chosenMentors;
 }
 function getMentorFileName(mentor: string): string {
  const model = mentor.toLowerCase().replace(/\s/g, '');
  return `mattw/${model}`;
 }
 async function getSystemPrompt(mentor: string, isLast: boolean, question: string): Promise<string> {
  ollama.setModel(getMentorFileName(mentor));
  const info = await ollama.showModelInfo()
  let SystemPrompt = info.system || '';
  SystemPrompt += ` You should continue the conversation as if you were ${mentor} and acknowledge the people before you in the conversation. You should adopt their mannerisms and tone, but also not use language they wouldn't use. If they are not known to know about the concept in the question, don't offer an answer. Your answer should be no longer than 1 paragraph. And definitely try not to sound like anyone else. Don't repeat any slang or phrases already used. And if it is a question the original ${mentor} wouldn't have know the answer to, just say that you don't know, in the style of ${mentor}. And think about the time the person lived. Don't use terminology that they wouldn't have used.`
  if (isLast) {
    SystemPrompt += ` End your answer with something like I hope our answers help you out`;
  } else {
    SystemPrompt += ` Remember, this is a conversation, so you don't need a conclusion, but end your answer with a question related to the first question: "${question}".`;
  }
  return SystemPrompt;
 }
 async function main() {
  const mentors = getMentors();
  const question = process.argv[2];
  let theConversation = `Here is the conversation so far.\nYou: ${question}\n`
  for await (const mentor of mentors) {
    const SystemPrompt = await getSystemPrompt(mentor, mentor === mentors[mentorCount - 1], question);
    ollama.setModel(getMentorFileName(mentor));
    ollama.setSystemPrompt(SystemPrompt);
    let output = '';
    process.stdout.write(`\n${mentor}: `);
    for await (const chunk of ollama.streamingGenerate(theConversation + `Continue the conversation as if you were ${mentor} on the question "${question}".`)) {
      if (chunk.response) {
        output += chunk.response;
        process.stdout.write(chunk.response);
      } else {
        process.stdout.write('\n');
      }
    }
    theConversation += `${mentor}: ${output}\n\n`
  }
 }
 main();
--- a/examples/typescript-mentors/package.json
+++ b/examples/typescript-mentors/package.json
@@ -0,0 +1,7 @@
 {
  "dependencies": {
    "fs": "^0.0.1-security",
    "ollama-node": "^0.0.3",
    "path": "^0.12.7"
  }
 }
--- a/format/bytes.go
+++ b/format/bytes.go
@@ -0,0 +1,16 @@
 package format
 import "fmt"
 func HumanBytes(b int64) string {
 	switch {
 	case b > 1000*1000*1000:
 		return fmt.Sprintf("%d GB", b/1000/1000/1000)
 	case b > 1000*1000:
 		return fmt.Sprintf("%d MB", b/1000/1000)
 	case b > 1000:
 		return fmt.Sprintf("%d KB", b/1000)
 	default:
 		return fmt.Sprintf("%d B", b)
 	}
 }
--- a/format/openssh.go
+++ b/format/openssh.go
@@ -0,0 +1,102 @@
 // Copyright 2012 The Go Authors. All rights reserved.
 // Use of this source code is governed by a BSD-style
 // license that can be found in the LICENSE file.
 // Code originally from https://go-review.googlesource.com/c/crypto/+/218620
 // TODO: replace with upstream once the above change is merged and released.
 package format
 import (
 	"crypto"
 	"crypto/ed25519"
 	"crypto/rand"
 	"encoding/binary"
 	"encoding/pem"
 	"fmt"
 	"golang.org/x/crypto/ssh"
 )
 const privateKeyAuthMagic = "openssh-key-v1\x00"
 type openSSHEncryptedPrivateKey struct {
 	CipherName string
 	KDFName    string
 	KDFOptions string
 	KeysCount  uint32
 	PubKey     []byte
 	KeyBlocks  []byte
 }
 type openSSHPrivateKey struct {
 	Check1  uint32
 	Check2  uint32
 	Keytype string
 	Rest    []byte `ssh:"rest"`
 }
 type openSSHEd25519PrivateKey struct {
 	Pub     []byte
 	Priv    []byte
 	Comment string
 	Pad     []byte `ssh:"rest"`
 }
 func OpenSSHPrivateKey(key crypto.PrivateKey, comment string) (*pem.Block, error) {
 	var check uint32
 	if err := binary.Read(rand.Reader, binary.BigEndian, &check); err != nil {
 		return nil, err
 	}
 	var pk1 openSSHPrivateKey
 	pk1.Check1 = check
 	pk1.Check2 = check
 	var w openSSHEncryptedPrivateKey
 	w.KeysCount = 1
 	if k, ok := key.(*ed25519.PrivateKey); ok {
 		key = *k
 	}
 	switch k := key.(type) {
 	case ed25519.PrivateKey:
 		pub, priv := k[32:], k
 		key := openSSHEd25519PrivateKey{
 			Pub:     pub,
 			Priv:    priv,
 			Comment: comment,
 		}
 		pk1.Keytype = ssh.KeyAlgoED25519
 		pk1.Rest = ssh.Marshal(key)
 		w.PubKey = ssh.Marshal(struct {
 			KeyType string
 			Pub     []byte
 		}{
 			ssh.KeyAlgoED25519, pub,
 		})
 	default:
 		return nil, fmt.Errorf("ssh: unknown key type %T", k)
 	}
 	w.KeyBlocks = openSSHPadding(ssh.Marshal(pk1), 8)
 	w.CipherName, w.KDFName, w.KDFOptions = "none", "none", ""
 	return &pem.Block{
 		Type:  "OPENSSH PRIVATE KEY",
 		Bytes: append([]byte(privateKeyAuthMagic), ssh.Marshal(w)...),
 	}, nil
 }
 func openSSHPadding(block []byte, blocksize int) []byte {
 	for i, j := 0, len(block); (j+i)%blocksize != 0; i++ {
 		block = append(block, byte(i+1))
 	}
 	return block
 }
--- a/format/time.go
+++ b/format/time.go
@@ -7,26 +7,14 @@ import (
 	"time"
 )
-// HumanDuration returns a human-readable approximation of a duration
+// humanDuration returns a human-readable approximation of a
-// (eg. "About a minute", "4 hours ago", etc.).
+// duration (eg. "About a minute", "4 hours ago", etc.).
-// Modified version of github.com/docker/go-units.HumanDuration
+func humanDuration(d time.Duration) string {
 func HumanDuration(d time.Duration) string {
 	return HumanDurationWithCase(d, true)
 }
 // HumanDurationWithCase returns a human-readable approximation of a
 // duration (eg. "About a minute", "4 hours ago", etc.). but allows
 // you to specify whether the first word should be capitalized
 // (eg. "About" vs. "about")
 func HumanDurationWithCase(d time.Duration, useCaps bool) string {
 	seconds := int(d.Seconds())
 	switch {
 	case seconds < 1:
-		if useCaps {
+		return "Less than a second"
 			return "Less than a second"
 		}
 		return "less than a second"
 	case seconds == 1:
 		return "1 second"
 	case seconds < 60:
@@ -36,10 +24,7 @@ func HumanDurationWithCase(d time.Duration, useCaps bool) string {
 	minutes := int(d.Minutes())
 	switch {
 	case minutes == 1:
-		if useCaps {
+		return "About a minute"
 			return "About a minute"
 		}
 		return "about a minute"
 	case minutes < 60:
 		return fmt.Sprintf("%d minutes", minutes)
 	}
@@ -47,10 +32,7 @@ func HumanDurationWithCase(d time.Duration, useCaps bool) string {
 	hours := int(math.Round(d.Hours()))
 	switch {
 	case hours == 1:
-		if useCaps {
+		return "About an hour"
 			return "About an hour"
 		}
 		return "about an hour"
 	case hours < 48:
 		return fmt.Sprintf("%d hours", hours)
 	case hours < 24*7*2:
@@ -65,77 +47,22 @@ func HumanDurationWithCase(d time.Duration, useCaps bool) string {
 }
 func HumanTime(t time.Time, zeroValue string) string {
-	return humanTimeWithCase(t, zeroValue, true)
+	return humanTime(t, zeroValue)
 }
 func HumanTimeLower(t time.Time, zeroValue string) string {
-	return humanTimeWithCase(t, zeroValue, false)
+	return strings.ToLower(humanTime(t, zeroValue))
 }
-func humanTimeWithCase(t time.Time, zeroValue string, useCaps bool) string {
+func humanTime(t time.Time, zeroValue string) string {
 	if t.IsZero() {
 		return zeroValue
 	}
 	delta := time.Since(t)
 	if delta < 0 {
-		return HumanDurationWithCase(-delta, useCaps) + " from now"
+		return humanDuration(-delta) + " from now"
 	}
-	return HumanDurationWithCase(delta, useCaps) + " ago"
+
-}
+	return humanDuration(delta) + " ago"
 // ExcatDuration returns a human readable hours/minutes/seconds or milliseconds format of a duration
 // the most precise level of duration is milliseconds
 func ExactDuration(d time.Duration) string {
 	if d.Seconds() < 1 {
 		if d.Milliseconds() == 1 {
 			return fmt.Sprintf("%d millisecond", d.Milliseconds())
 		}
 		return fmt.Sprintf("%d milliseconds", d.Milliseconds())
 	}
 	var readableDur strings.Builder
 	dur := d.String()
 	// split the default duration string format of 0h0m0s into something nicer to read
 	h := strings.Split(dur, "h")
 	if len(h) > 1 {
 		hours := h[0]
 		if hours == "1" {
 			readableDur.WriteString(fmt.Sprintf("%s hour ", hours))
 		} else {
 			readableDur.WriteString(fmt.Sprintf("%s hours ", hours))
 		}
 		dur = h[1]
 	}
 	m := strings.Split(dur, "m")
 	if len(m) > 1 {
 		mins := m[0]
 		switch mins {
 		case "0":
 			// skip
 		case "1":
 			readableDur.WriteString(fmt.Sprintf("%s minute ", mins))
 		default:
 			readableDur.WriteString(fmt.Sprintf("%s minutes ", mins))
 		}
 		dur = m[1]
 	}
 	s := strings.Split(dur, "s")
 	if len(s) > 0 {
 		sec := s[0]
 		switch sec {
 		case "0":
 			// skip
 		case "1":
 			readableDur.WriteString(fmt.Sprintf("%s second ", sec))
 		default:
 			readableDur.WriteString(fmt.Sprintf("%s seconds ", sec))
 		}
 	}
 	return strings.TrimSpace(readableDur.String())
 }
--- a/format/time_test.go
+++ b/format/time_test.go
@@ -11,92 +11,25 @@ func assertEqual(t *testing.T, a interface{}, b interface{}) {
 	}
 }
 func TestHumanDuration(t *testing.T) {
 	day := 24 * time.Hour
 	week := 7 * day
 	month := 30 * day
 	year := 365 * day
 	assertEqual(t, "Less than a second", HumanDuration(450*time.Millisecond))
 	assertEqual(t, "Less than a second", HumanDurationWithCase(450*time.Millisecond, true))
 	assertEqual(t, "less than a second", HumanDurationWithCase(450*time.Millisecond, false))
 	assertEqual(t, "1 second", HumanDuration(1*time.Second))
 	assertEqual(t, "45 seconds", HumanDuration(45*time.Second))
 	assertEqual(t, "46 seconds", HumanDuration(46*time.Second))
 	assertEqual(t, "59 seconds", HumanDuration(59*time.Second))
 	assertEqual(t, "About a minute", HumanDuration(60*time.Second))
 	assertEqual(t, "About a minute", HumanDurationWithCase(1*time.Minute, true))
 	assertEqual(t, "about a minute", HumanDurationWithCase(1*time.Minute, false))
 	assertEqual(t, "3 minutes", HumanDuration(3*time.Minute))
 	assertEqual(t, "35 minutes", HumanDuration(35*time.Minute))
 	assertEqual(t, "35 minutes", HumanDuration(35*time.Minute+40*time.Second))
 	assertEqual(t, "45 minutes", HumanDuration(45*time.Minute))
 	assertEqual(t, "45 minutes", HumanDuration(45*time.Minute+40*time.Second))
 	assertEqual(t, "46 minutes", HumanDuration(46*time.Minute))
 	assertEqual(t, "59 minutes", HumanDuration(59*time.Minute))
 	assertEqual(t, "About an hour", HumanDuration(1*time.Hour))
 	assertEqual(t, "About an hour", HumanDurationWithCase(1*time.Hour+29*time.Minute, true))
 	assertEqual(t, "about an hour", HumanDurationWithCase(1*time.Hour+29*time.Minute, false))
 	assertEqual(t, "2 hours", HumanDuration(1*time.Hour+31*time.Minute))
 	assertEqual(t, "2 hours", HumanDuration(1*time.Hour+59*time.Minute))
 	assertEqual(t, "3 hours", HumanDuration(3*time.Hour))
 	assertEqual(t, "3 hours", HumanDuration(3*time.Hour+29*time.Minute))
 	assertEqual(t, "4 hours", HumanDuration(3*time.Hour+31*time.Minute))
 	assertEqual(t, "4 hours", HumanDuration(3*time.Hour+59*time.Minute))
 	assertEqual(t, "4 hours", HumanDuration(3*time.Hour+60*time.Minute))
 	assertEqual(t, "24 hours", HumanDuration(24*time.Hour))
 	assertEqual(t, "36 hours", HumanDuration(1*day+12*time.Hour))
 	assertEqual(t, "2 days", HumanDuration(2*day))
 	assertEqual(t, "7 days", HumanDuration(7*day))
 	assertEqual(t, "13 days", HumanDuration(13*day+5*time.Hour))
 	assertEqual(t, "2 weeks", HumanDuration(2*week))
 	assertEqual(t, "2 weeks", HumanDuration(2*week+4*day))
 	assertEqual(t, "3 weeks", HumanDuration(3*week))
 	assertEqual(t, "4 weeks", HumanDuration(4*week))
 	assertEqual(t, "4 weeks", HumanDuration(4*week+3*day))
 	assertEqual(t, "4 weeks", HumanDuration(1*month))
 	assertEqual(t, "6 weeks", HumanDuration(1*month+2*week))
 	assertEqual(t, "2 months", HumanDuration(2*month))
 	assertEqual(t, "2 months", HumanDuration(2*month+2*week))
 	assertEqual(t, "3 months", HumanDuration(3*month))
 	assertEqual(t, "3 months", HumanDuration(3*month+1*week))
 	assertEqual(t, "5 months", HumanDuration(5*month+2*week))
 	assertEqual(t, "13 months", HumanDuration(13*month))
 	assertEqual(t, "23 months", HumanDuration(23*month))
 	assertEqual(t, "24 months", HumanDuration(24*month))
 	assertEqual(t, "2 years", HumanDuration(24*month+2*week))
 	assertEqual(t, "3 years", HumanDuration(3*year+2*month))
 }
 func TestHumanTime(t *testing.T) {
 	now := time.Now()
 	t.Run("zero value", func(t *testing.T) {
 		assertEqual(t, HumanTime(time.Time{}, "never"), "never")
 	})
 	t.Run("time in the future", func(t *testing.T) {
 		v := now.Add(48 * time.Hour)
 		assertEqual(t, HumanTime(v, ""), "2 days from now")
 	})
 	t.Run("time in the past", func(t *testing.T) {
 		v := now.Add(-48 * time.Hour)
 		assertEqual(t, HumanTime(v, ""), "2 days ago")
 	})
 }
-func TestExactDuration(t *testing.T) {
+	t.Run("soon", func(t *testing.T) {
-	assertEqual(t, "1 millisecond", ExactDuration(1*time.Millisecond))
+		v := now.Add(800*time.Millisecond)
-	assertEqual(t, "10 milliseconds", ExactDuration(10*time.Millisecond))
+		assertEqual(t, HumanTime(v, ""), "Less than a second from now")
-	assertEqual(t, "1 second", ExactDuration(1*time.Second))
+	})
 	assertEqual(t, "10 seconds", ExactDuration(10*time.Second))
 	assertEqual(t, "1 minute", ExactDuration(1*time.Minute))
 	assertEqual(t, "10 minutes", ExactDuration(10*time.Minute))
 	assertEqual(t, "1 hour", ExactDuration(1*time.Hour))
 	assertEqual(t, "10 hours", ExactDuration(10*time.Hour))
 	assertEqual(t, "1 hour 1 second", ExactDuration(1*time.Hour+1*time.Second))
 	assertEqual(t, "1 hour 10 seconds", ExactDuration(1*time.Hour+10*time.Second))
 	assertEqual(t, "1 hour 1 minute", ExactDuration(1*time.Hour+1*time.Minute))
 	assertEqual(t, "1 hour 10 minutes", ExactDuration(1*time.Hour+10*time.Minute))
 	assertEqual(t, "1 hour 1 minute 1 second", ExactDuration(1*time.Hour+1*time.Minute+1*time.Second))
 	assertEqual(t, "10 hours 10 minutes 10 seconds", ExactDuration(10*time.Hour+10*time.Minute+10*time.Second))
 }
--- a/ggml-metal.metal
+++ b/ggml-metal.metal
@@ -1 +0,0 @@
 llama/ggml-metal.metal
--- a/go.mod
+++ b/go.mod
@@ -8,16 +8,16 @@ require (
 	github.com/mattn/go-runewidth v0.0.14
 	github.com/mitchellh/colorstring v0.0.0-20190213212951-d06e56a500db
 	github.com/olekukonko/tablewriter v0.0.5
 	github.com/pdevine/readline v1.5.2
 	github.com/spf13/cobra v1.7.0
 	golang.org/x/sync v0.3.0
 )
 require github.com/rivo/uniseg v0.2.0 // indirect
 require (
 	dario.cat/mergo v1.0.0
 	github.com/bytedance/sonic v1.9.1 // indirect
 	github.com/chenzhuoyu/base64x v0.0.0-20221115062448-fe3a3abad311 // indirect
 	github.com/chzyer/readline v1.5.1
 	github.com/gabriel-vasile/mimetype v1.4.2 // indirect
 	github.com/gin-contrib/cors v1.4.0
 	github.com/gin-contrib/sse v0.1.0 // indirect
@@ -33,16 +33,19 @@ require (
 	github.com/mattn/go-isatty v0.0.19 // indirect
 	github.com/modern-go/concurrent v0.0.0-20180306012644-bacd9c7ef1dd // indirect
 	github.com/modern-go/reflect2 v1.0.2 // indirect
 	github.com/pbnjay/memory v0.0.0-20210728143218-7b4eea64cf58
 	github.com/pelletier/go-toml/v2 v2.0.8 // indirect
 	github.com/spf13/pflag v1.0.5 // indirect
 	github.com/twitchyliquid64/golang-asm v0.15.1 // indirect
 	github.com/ugorji/go/codec v1.2.11 // indirect
 	golang.org/x/arch v0.3.0 // indirect
-	golang.org/x/crypto v0.10.0 // indirect
+	golang.org/x/crypto v0.10.0
 	golang.org/x/exp v0.0.0-20230817173708-d852ddb80c63
 	golang.org/x/net v0.10.0 // indirect
-	golang.org/x/sys v0.10.0 // indirect
+	golang.org/x/sys v0.11.0 // indirect
 	golang.org/x/term v0.10.0
 	golang.org/x/text v0.10.0 // indirect
 	gonum.org/v1/gonum v0.13.0
 	google.golang.org/protobuf v1.30.0 // indirect
 	gopkg.in/yaml.v3 v3.0.1 // indirect
 )
--- a/go.sum
+++ b/go.sum
@@ -1,5 +1,3 @@
 dario.cat/mergo v1.0.0 h1:AGCNq9Evsj31mOgNPcLyXc+4PNABt905YmuqPYYpBWk=
 dario.cat/mergo v1.0.0/go.mod h1:uNxQE+84aUszobStD9th8a29P2fMDhsBdgRYvZOxGmk=
 github.com/bytedance/sonic v1.5.0/go.mod h1:ED5hyg4y6t3/9Ku1R6dU/4KyJ48DZ4jPhfY1O2AihPM=
 github.com/bytedance/sonic v1.9.1 h1:6iJ6NqdoxCDr6mbY8h18oSO+cShGSMRGCEo7F2h0x8s=
 github.com/bytedance/sonic v1.9.1/go.mod h1:i736AoUSYt75HyZLoJW9ERYxcy6eaN6h4BZXU064P/U=
@@ -8,8 +6,6 @@ github.com/chenzhuoyu/base64x v0.0.0-20221115062448-fe3a3abad311 h1:qSGYFH7+jGhD
 github.com/chenzhuoyu/base64x v0.0.0-20221115062448-fe3a3abad311/go.mod h1:b583jCggY9gE99b6G5LEC39OIiVsWj+R97kbl5odCEk=
 github.com/chzyer/logex v1.2.1 h1:XHDu3E6q+gdHgsdTPH6ImJMIp436vR6MPtH8gP05QzM=
 github.com/chzyer/logex v1.2.1/go.mod h1:JLbx6lG2kDbNRFnfkgvh4eRJRPX1QCoOIWomwysCBrQ=
 github.com/chzyer/readline v1.5.1 h1:upd/6fQk4src78LMRzh5vItIt361/o4uq553V8B5sGI=
 github.com/chzyer/readline v1.5.1/go.mod h1:Eh+b79XXUwfKfcPLepksvw2tcLE/Ct21YObkaSkeBlk=
 github.com/chzyer/test v1.0.0 h1:p3BQDXSxOhOG0P9z6/hGnII4LGiEPOYBhs8asl/fC04=
 github.com/chzyer/test v1.0.0/go.mod h1:2JlltgoNkt4TW/z9V/IzDdFaMTM2JPIi26O1pF38GC8=
 github.com/cpuguy83/go-md2man/v2 v2.0.2/go.mod h1:tgQtvFlXSQOSOSIRvRPT7W67SCa46tRHOmNcaadrF8o=
@@ -80,6 +76,10 @@ github.com/modern-go/reflect2 v1.0.2 h1:xBagoLtFs94CBntxluKeaWgTMpvLxC4ur3nMaC9G
 github.com/modern-go/reflect2 v1.0.2/go.mod h1:yWuevngMOJpCy52FWWMvUC8ws7m/LJsjYzDa0/r8luk=
 github.com/olekukonko/tablewriter v0.0.5 h1:P2Ga83D34wi1o9J6Wh1mRuqd4mF/x/lgBS7N7AbDhec=
 github.com/olekukonko/tablewriter v0.0.5/go.mod h1:hPp6KlRPjbx+hW8ykQs1w3UBbZlj6HuIJcUGPhkA7kY=
 github.com/pbnjay/memory v0.0.0-20210728143218-7b4eea64cf58 h1:onHthvaw9LFnH4t2DcNVpwGmV9E1BkGknEliJkfwQj0=
 github.com/pbnjay/memory v0.0.0-20210728143218-7b4eea64cf58/go.mod h1:DXv8WO4yhMYhSNPKjeNKa5WY9YCIEBRbNzFFPJbWO6Y=
 github.com/pdevine/readline v1.5.2 h1:oz6Y5GdTmhPG+08hhxcAvtHitSANWuA2100Sppb38xI=
 github.com/pdevine/readline v1.5.2/go.mod h1:na/LbuE5PYwxI7GyopWdIs3U8HVe89lYlNTFTXH3wOw=
 github.com/pelletier/go-toml/v2 v2.0.1/go.mod h1:r9LEWfGN8R5k0VXJ+0BkIe7MYkRdwZOjgMj2KwnJFUo=
 github.com/pelletier/go-toml/v2 v2.0.8 h1:0ctb6s9mE31h0/lhu+J6OPmVeDxJn+kYnJc2jZR9tGQ=
 github.com/pelletier/go-toml/v2 v2.0.8/go.mod h1:vuYfssBdrU2XDZ9bYydBu6t+6a6PYNcZljzZR9VXg+4=
@@ -120,9 +120,13 @@ golang.org/x/arch v0.3.0/go.mod h1:5om86z9Hs0C8fWVUuoMHwpExlXzs5Tkyp9hOrfG7pp8=
 golang.org/x/crypto v0.0.0-20210711020723-a769d52b0f97/go.mod h1:GvvjBRRGRdwPK5ydBHafDWAxML/pGHZbMvKqRZ5+Abc=
 golang.org/x/crypto v0.10.0 h1:LKqV2xt9+kDzSTfOhx4FrkEBcMrAgHSYgzywV9zcGmM=
 golang.org/x/crypto v0.10.0/go.mod h1:o4eNf7Ede1fv+hwOwZsTHl9EsPFO6q6ZvYR8vYfY45I=
 golang.org/x/exp v0.0.0-20230817173708-d852ddb80c63 h1:m64FZMko/V45gv0bNmrNYoDEq8U5YUhetc9cBWKS1TQ=
 golang.org/x/exp v0.0.0-20230817173708-d852ddb80c63/go.mod h1:0v4NqG35kSWCMzLaMeX+IQrlSnVE/bqGSyC2cz/9Le8=
 golang.org/x/net v0.0.0-20210226172049-e18ecbb05110/go.mod h1:m0MpNAwzfU5UDzcl9v0D8zg8gWTRqZa9RBIspLL5mdg=
 golang.org/x/net v0.10.0 h1:X2//UzNDwYmtCLn7To6G58Wr6f5ahEAQgKNzv9Y951M=
 golang.org/x/net v0.10.0/go.mod h1:0qNGK6F8kojg2nk9dLZ2mShWaEBan6FAoqfSigmmuDg=
 golang.org/x/sync v0.3.0 h1:ftCYgMx6zT/asHUrPw8BLLscYtGznsLAnjq5RH9P66E=
 golang.org/x/sync v0.3.0/go.mod h1:FU7BRWz2tNW+3quACPkgCx/L+uEAv1htQ0V83Z9Rj+Y=
 golang.org/x/sys v0.0.0-20201119102817-f84b799fce68/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
 golang.org/x/sys v0.0.0-20210615035016-665e8c7367d1/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
 golang.org/x/sys v0.0.0-20210630005230-0f9fa26af87c/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
@@ -130,8 +134,8 @@ golang.org/x/sys v0.0.0-20210806184541-e5e7981a1069/go.mod h1:oPkhp1MJrh7nUepCBc
 golang.org/x/sys v0.0.0-20220310020820-b874c991c1a5/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
 golang.org/x/sys v0.0.0-20220704084225-05e143d24a9e/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
 golang.org/x/sys v0.6.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
-golang.org/x/sys v0.10.0 h1:SqMFp9UcQJZa+pmYuAKjd9xq1f0j5rLcDIk0mj4qAsA=
+golang.org/x/sys v0.11.0 h1:eG7RXZHdqOJ1i+0lgLgCpSXAp6M3LYlAo6osgSi0xOM=
-golang.org/x/sys v0.10.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
+golang.org/x/sys v0.11.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
 golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
 golang.org/x/term v0.10.0 h1:3R7pNqamzBraeqj/Tj8qt1aQ2HpmlC+Cx/qL/7hn4/c=
 golang.org/x/term v0.10.0/go.mod h1:lpqdcUyK/oCiQxvxVrppt5ggO2KCZ5QblwqPnfZ6d5o=
@@ -141,6 +145,8 @@ golang.org/x/text v0.10.0 h1:UpjohKhiEgNc0CSauXmwYftY1+LlaC75SJwh0SgCX58=
 golang.org/x/text v0.10.0/go.mod h1:TvPlkZtksWOMsz7fbANvkp4WM8x/WCo/om8BMLbz+aE=
 golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ=
 golang.org/x/xerrors v0.0.0-20191204190536-9bdfabe68543/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0=
 gonum.org/v1/gonum v0.13.0 h1:a0T3bh+7fhRyqeNbiC3qVHYmkiQgit3wnNan/2c0HMM=
 gonum.org/v1/gonum v0.13.0/go.mod h1:/WPYRckkfWrhWefxyYTfrTtQR0KH4iyHNuzxqXAKyAU=
 google.golang.org/protobuf v1.26.0-rc.1/go.mod h1:jlhhOSvTdKEhbULTjvd4ARK9grFBp09yW+WbY/TyQbw=
 google.golang.org/protobuf v1.28.0/go.mod h1:HV8QOd/L58Z+nl8r43ehVNZIU/HEI6OcFqwMG9pJV4I=
 google.golang.org/protobuf v1.30.0 h1:kPPoIgf3TsEvrm0PFe15JQ+570QVxYzEvvHqChK+cng=
--- a/library/.gitignore
+++ b/library/.gitignore
@@ -1 +0,0 @@
 models
--- a/library/downloads
+++ b/library/downloads
@@ -1,7 +0,0 @@
 https://huggingface.co/TheBloke/orca_mini_3B-GGML/resolve/main/orca-mini-3b.ggmlv3.q4_0.bin e84705205f71dd55be7b24a778f248f0eda9999a125d313358c087e092d83148
 https://huggingface.co/TheBloke/Nous-Hermes-13B-GGML/resolve/main/nous-hermes-13b.ggmlv3.q4_0.bin d1735b93e1dc503f1045ccd6c8bd73277b18ba892befd1dc29e9b9a7822ed998
 https://huggingface.co/TheBloke/vicuna-7B-v1.3-GGML/resolve/main/vicuna-7b-v1.3.ggmlv3.q4_0.bin 23ce5ed290b56a19305178b9ada2c3d96036bd69a6c18304b6158eb6672d6c0f
 https://huggingface.co/TheBloke/Wizard-Vicuna-13B-Uncensored-GGML/resolve/main/Wizard-Vicuna-13B-Uncensored.ggmlv3.q4_0.bin 1f08b147a5bce41cfcbb3fd5d51ba765dea1786e15b5655ab69ba3a337a893b7
 https://huggingface.co/TheBloke/Llama-2-7B-GGML/resolve/main/llama-2-7b.ggmlv3.q4_0.bin bfa26d855e44629c4cf919985e90bd7fa03b77eea1676791519e39a4d45fd4d5
 https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGML/resolve/main/llama-2-7b-chat.ggmlv3.q4_0.bin 8daa9615cce30c259a9555b1cc250d461d1bc69980a274b44d7eda0be78076d8
 https://huggingface.co/TheBloke/Llama-2-13B-chat-GGML/resolve/main/llama-2-13b-chat.ggmlv3.q4_0.bin f79142715bc9539a2edbb4b253548db8b34fac22736593eeaa28555874476e30
--- a/library/modelfiles/llama2
+++ b/library/modelfiles/llama2
@@ -1,147 +0,0 @@
 FROM ../models/llama-2-7b-chat.ggmlv3.q4_0.bin
 TEMPLATE """
 {{- if .First }}
 <<SYS>>
 {{ .System }}
 <</SYS>>
 {{- end }}
 [INST] {{ .Prompt }} [/INST]
 """
 SYSTEM """
 You are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature.
 If a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information.
 """
 LICENSE """
 Llama 2 Community License Agreement
 Llama 2 Version Release Date: July 18, 2023
 “Agreement” means the terms and conditions for use, reproduction, distribution and modification of the Llama Materials set forth herein.
 “Documentation” means the specifications, manuals and documentation accompanying Llama 2 distributed by Meta at ai.meta.com/resources/models-and-libraries/llama-downloads/.
 “Licensee” or “you” means you, or your employer or any other person or entity (if you are entering into this Agreement on such person or entity’s behalf), of the age required under applicable laws, rules or regulations to provide legal consent and that has legal authority to bind your employer or such other person or entity if you are entering in this Agreement on their behalf.
 “Llama 2” means the foundational large language models and software and algorithms, including machine-learning model code, trained model weights, inference-enabling code, training-enabling code, fine-tuning enabling code and other elements of the foregoing distributed by Meta at ai.meta.com/resources/models-and-libraries/llama-downloads/.
 “Llama Materials” means, collectively, Meta’s proprietary Llama 2 and Documentation (and any portion thereof) made available under this Agreement.
 “Meta” or “we” means Meta Platforms Ireland Limited (if you are located in or, if you are an entity, your principal place of business is in the EEA or Switzerland) and Meta Platforms, Inc. (if you are located outside of the EEA or Switzerland).
 By clicking “I Accept” below or by using or distributing any portion or element of the Llama Materials, you agree to be bound by this Agreement.
 1. License Rights and Redistribution.
 a. Grant of Rights. You are granted a non-exclusive, worldwide, non-transferable and royalty-free limited license under Meta’s intellectual property or other rights owned by Meta embodied in the Llama Materials to use, reproduce, distribute, copy, create derivative works of, and make modifications to the Llama Materials.
 b. Redistribution and Use.
 i. If you distribute or make the Llama Materials, or any derivative works thereof, available to a third party, you shall provide a copy of this Agreement to such third party.
 ii. If you receive Llama Materials, or any derivative works thereof, from a Licensee as part of an integrated end user product, then Section 2 of this Agreement will not apply to you.
 iii. You must retain in all copies of the Llama Materials that you distribute the following attribution notice within a “Notice” text file distributed as a part of such copies: “Llama 2 is licensed under the LLAMA 2 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.”
 iv. Your use of the Llama Materials must comply with applicable laws and regulations (including trade compliance laws and regulations) and adhere to the Acceptable Use Policy for the Llama Materials (available at https://ai.meta.com/llama/use-policy), which is hereby incorporated by reference into this Agreement.
 v. You will not use the Llama Materials or any output or results of the Llama Materials to improve any other large language model (excluding Llama 2 or derivative works thereof).
 2. Additional Commercial Terms. If, on the Llama 2 version release date, the monthly active users of the products or services made available by or for Licensee, or Licensee’s affiliates, is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta, which Meta may grant to you in its sole discretion, and you are not authorized to exercise any of the rights under this Agreement unless or until Meta otherwise expressly grants you such rights.
 3. Disclaimer of Warranty. UNLESS REQUIRED BY APPLICABLE LAW, THE LLAMA MATERIALS AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED ON AN “AS IS” BASIS, WITHOUT WARRANTIES OF ANY KIND, EITHER EXPRESS OR IMPLIED, INCLUDING, WITHOUT LIMITATION, ANY WARRANTIES OF TITLE, NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE. YOU ARE SOLELY RESPONSIBLE FOR DETERMINING THE APPROPRIATENESS OF USING OR REDISTRIBUTING THE LLAMA MATERIALS AND ASSUME ANY RISKS ASSOCIATED WITH YOUR USE OF THE LLAMA MATERIALS AND ANY OUTPUT AND RESULTS.
 4. Limitation of Liability. IN NO EVENT WILL META OR ITS AFFILIATES BE LIABLE UNDER ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, TORT, NEGLIGENCE, PRODUCTS LIABILITY, OR OTHERWISE, ARISING OUT OF THIS AGREEMENT, FOR ANY LOST PROFITS OR ANY INDIRECT, SPECIAL, CONSEQUENTIAL, INCIDENTAL, EXEMPLARY OR PUNITIVE DAMAGES, EVEN IF META OR ITS AFFILIATES HAVE BEEN ADVISED OF THE POSSIBILITY OF ANY OF THE FOREGOING.
 5. Intellectual Property.
 a. No trademark licenses are granted under this Agreement, and in connection with the Llama Materials, neither Meta nor Licensee may use any name or mark owned by or associated with the other or any of its affiliates, except as required for reasonable and customary use in describing and redistributing the Llama Materials.
 b. Subject to Meta’s ownership of Llama Materials and derivatives made by or for Meta, with respect to any derivative works and modifications of the Llama Materials that are made by you, as between you and Meta, you are and will be the owner of such derivative works and modifications.
 c. If you institute litigation or other proceedings against Meta or any entity (including a cross-claim or counterclaim in a lawsuit) alleging that the Llama Materials or Llama 2 outputs or results, or any portion of any of the foregoing, constitutes infringement of intellectual property or other rights owned or licensable by you, then any licenses granted to you under this Agreement shall terminate as of the date such litigation or claim is filed or instituted. You will indemnify and hold harmless Meta from and against any claim by any third party arising out of or related to your use or distribution of the Llama Materials.
 6. Term and Termination. The term of this Agreement will commence upon your acceptance of this Agreement or access to the Llama Materials and will continue in full force and effect until terminated in accordance with the terms and conditions herein. Meta may terminate this Agreement if you are in breach of any term or condition of this Agreement. Upon termination of this Agreement, you shall delete and cease use of the Llama Materials. Sections 3, 4 and 7 shall survive the termination of this Agreement.
 7. Governing Law and Jurisdiction. This Agreement will be governed and construed under the laws of the State of California without regard to choice of law principles, and the UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement. The courts of California shall have exclusive jurisdiction of any dispute arising out of this Agreement.
 """
 LICENSE """
 Llama 2 Acceptable Use Policy
 Meta is committed to promoting safe and fair use of its tools and features, including Llama 2. If you access or use Llama 2, you agree to this Acceptable Use Policy (“Policy”). The most recent copy of this policy can be found at ai.meta.com/llama/use-policy.
 Prohibited Uses
 We want everyone to use Llama 2 safely and responsibly. You agree you will not use, or allow others to use, Llama 2 to:
 1. Violate the law or others’ rights, including to:
 a. Engage in, promote, generate, contribute to, encourage, plan, incite, or further illegal or unlawful activity or content, such as:
 i. Violence or terrorism
 ii. Exploitation or harm to children, including the solicitation, creation, acquisition, or dissemination of child exploitative content or failure to report Child Sexual Abuse Material
 b. Human trafficking, exploitation, and sexual violence
 iii. The illegal distribution of information or materials to minors, including obscene materials, or failure to employ legally required age-gating in connection with such information or materials.
 iv. Sexual solicitation
 vi. Any other criminal activity
 c. Engage in, promote, incite, or facilitate the harassment, abuse, threatening, or bullying of individuals or groups of individuals
 d. Engage in, promote, incite, or facilitate discrimination or other unlawful or harmful conduct in the provision of employment, employment benefits, credit, housing, other economic benefits, or other essential goods and services
 e. Engage in the unauthorized or unlicensed practice of any profession including, but not limited to, financial, legal, medical/health, or related professional practices
 f. Collect, process, disclose, generate, or infer health, demographic, or other sensitive personal or private information about individuals without rights and consents required by applicable laws
 g. Engage in or facilitate any action or generate any content that infringes, misappropriates, or otherwise violates any third-party rights, including the outputs or results of any products or services using the Llama 2 Materials
 h. Create, generate, or facilitate the creation of malicious code, malware, computer viruses or do anything else that could disable, overburden, interfere with or impair the proper working, integrity, operation or appearance of a website or computer system
 2. Engage in, promote, incite, facilitate, or assist in the planning or development of activities that present a risk of death or bodily harm to individuals, including use of Llama 2 related to the following:
 a. Military, warfare, nuclear industries or applications, espionage, use for materials or activities that are subject to the International Traffic Arms Regulations (ITAR) maintained by the United States Department of State
 b. Guns and illegal weapons (including weapon development)
 c. Illegal drugs and regulated/controlled substances
 d. Operation of critical infrastructure, transportation technologies, or heavy machinery
 e. Self-harm or harm to others, including suicide, cutting, and eating disorders
 f. Any content intended to incite or promote violence, abuse, or any infliction of bodily harm to an individual
 3. Intentionally deceive or mislead others, including use of Llama 2 related to the following:
 a. Generating, promoting, or furthering fraud or the creation or promotion of disinformation
 b. Generating, promoting, or furthering defamatory content, including the creation of defamatory statements, images, or other content
 c. Generating, promoting, or further distributing spam
 d. Impersonating another individual without consent, authorization, or legal right
 e. Representing that the use of Llama 2 or outputs are human-generated
 f. Generating or facilitating false online engagement, including fake reviews and other means of fake online engagement
 4. Fail to appropriately disclose to end users any known dangers of your AI system
 Please report any violation of this Policy, software “bug,” or other problems that could lead to a violation of this Policy through one of the following means:
 Reporting issues with the model: github.com/facebookresearch/llama
 Reporting risky content generated by the model: developers.facebook.com/llama_output_feedback
 Reporting bugs and security concerns: facebook.com/whitehat/info
 Reporting violations of the Acceptable Use Policy or unlicensed uses of Llama: LlamaUseReport@meta.com
 """
--- a/library/modelfiles/llama2_13b
+++ b/library/modelfiles/llama2_13b
@@ -1,147 +0,0 @@
 FROM ../models/llama-2-13b-chat.ggmlv3.q4_0.bin
 TEMPLATE """
 {{- if .First }}
 <<SYS>>
 {{ .System }}
 <</SYS>>
 {{- end }}
 [INST] {{ .Prompt }} [/INST]
 """
 SYSTEM """
 You are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature.
 If a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information.
 """
 LICENSE """
 Llama 2 Community License Agreement
 Llama 2 Version Release Date: July 18, 2023
 “Agreement” means the terms and conditions for use, reproduction, distribution and modification of the Llama Materials set forth herein.
 “Documentation” means the specifications, manuals and documentation accompanying Llama 2 distributed by Meta at ai.meta.com/resources/models-and-libraries/llama-downloads/.
 “Licensee” or “you” means you, or your employer or any other person or entity (if you are entering into this Agreement on such person or entity’s behalf), of the age required under applicable laws, rules or regulations to provide legal consent and that has legal authority to bind your employer or such other person or entity if you are entering in this Agreement on their behalf.
 “Llama 2” means the foundational large language models and software and algorithms, including machine-learning model code, trained model weights, inference-enabling code, training-enabling code, fine-tuning enabling code and other elements of the foregoing distributed by Meta at ai.meta.com/resources/models-and-libraries/llama-downloads/.
 “Llama Materials” means, collectively, Meta’s proprietary Llama 2 and Documentation (and any portion thereof) made available under this Agreement.
 “Meta” or “we” means Meta Platforms Ireland Limited (if you are located in or, if you are an entity, your principal place of business is in the EEA or Switzerland) and Meta Platforms, Inc. (if you are located outside of the EEA or Switzerland).
 By clicking “I Accept” below or by using or distributing any portion or element of the Llama Materials, you agree to be bound by this Agreement.
 1. License Rights and Redistribution.
 a. Grant of Rights. You are granted a non-exclusive, worldwide, non-transferable and royalty-free limited license under Meta’s intellectual property or other rights owned by Meta embodied in the Llama Materials to use, reproduce, distribute, copy, create derivative works of, and make modifications to the Llama Materials.
 b. Redistribution and Use.
 i. If you distribute or make the Llama Materials, or any derivative works thereof, available to a third party, you shall provide a copy of this Agreement to such third party.
 ii. If you receive Llama Materials, or any derivative works thereof, from a Licensee as part of an integrated end user product, then Section 2 of this Agreement will not apply to you.
 iii. You must retain in all copies of the Llama Materials that you distribute the following attribution notice within a “Notice” text file distributed as a part of such copies: “Llama 2 is licensed under the LLAMA 2 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.”
 iv. Your use of the Llama Materials must comply with applicable laws and regulations (including trade compliance laws and regulations) and adhere to the Acceptable Use Policy for the Llama Materials (available at https://ai.meta.com/llama/use-policy), which is hereby incorporated by reference into this Agreement.
 v. You will not use the Llama Materials or any output or results of the Llama Materials to improve any other large language model (excluding Llama 2 or derivative works thereof).
 2. Additional Commercial Terms. If, on the Llama 2 version release date, the monthly active users of the products or services made available by or for Licensee, or Licensee’s affiliates, is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta, which Meta may grant to you in its sole discretion, and you are not authorized to exercise any of the rights under this Agreement unless or until Meta otherwise expressly grants you such rights.
 3. Disclaimer of Warranty. UNLESS REQUIRED BY APPLICABLE LAW, THE LLAMA MATERIALS AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED ON AN “AS IS” BASIS, WITHOUT WARRANTIES OF ANY KIND, EITHER EXPRESS OR IMPLIED, INCLUDING, WITHOUT LIMITATION, ANY WARRANTIES OF TITLE, NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE. YOU ARE SOLELY RESPONSIBLE FOR DETERMINING THE APPROPRIATENESS OF USING OR REDISTRIBUTING THE LLAMA MATERIALS AND ASSUME ANY RISKS ASSOCIATED WITH YOUR USE OF THE LLAMA MATERIALS AND ANY OUTPUT AND RESULTS.
 4. Limitation of Liability. IN NO EVENT WILL META OR ITS AFFILIATES BE LIABLE UNDER ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, TORT, NEGLIGENCE, PRODUCTS LIABILITY, OR OTHERWISE, ARISING OUT OF THIS AGREEMENT, FOR ANY LOST PROFITS OR ANY INDIRECT, SPECIAL, CONSEQUENTIAL, INCIDENTAL, EXEMPLARY OR PUNITIVE DAMAGES, EVEN IF META OR ITS AFFILIATES HAVE BEEN ADVISED OF THE POSSIBILITY OF ANY OF THE FOREGOING.
 5. Intellectual Property.
 a. No trademark licenses are granted under this Agreement, and in connection with the Llama Materials, neither Meta nor Licensee may use any name or mark owned by or associated with the other or any of its affiliates, except as required for reasonable and customary use in describing and redistributing the Llama Materials.
 b. Subject to Meta’s ownership of Llama Materials and derivatives made by or for Meta, with respect to any derivative works and modifications of the Llama Materials that are made by you, as between you and Meta, you are and will be the owner of such derivative works and modifications.
 c. If you institute litigation or other proceedings against Meta or any entity (including a cross-claim or counterclaim in a lawsuit) alleging that the Llama Materials or Llama 2 outputs or results, or any portion of any of the foregoing, constitutes infringement of intellectual property or other rights owned or licensable by you, then any licenses granted to you under this Agreement shall terminate as of the date such litigation or claim is filed or instituted. You will indemnify and hold harmless Meta from and against any claim by any third party arising out of or related to your use or distribution of the Llama Materials.
 6. Term and Termination. The term of this Agreement will commence upon your acceptance of this Agreement or access to the Llama Materials and will continue in full force and effect until terminated in accordance with the terms and conditions herein. Meta may terminate this Agreement if you are in breach of any term or condition of this Agreement. Upon termination of this Agreement, you shall delete and cease use of the Llama Materials. Sections 3, 4 and 7 shall survive the termination of this Agreement.
 7. Governing Law and Jurisdiction. This Agreement will be governed and construed under the laws of the State of California without regard to choice of law principles, and the UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement. The courts of California shall have exclusive jurisdiction of any dispute arising out of this Agreement.
 """
 LICENSE """
 Llama 2 Acceptable Use Policy
 Meta is committed to promoting safe and fair use of its tools and features, including Llama 2. If you access or use Llama 2, you agree to this Acceptable Use Policy (“Policy”). The most recent copy of this policy can be found at ai.meta.com/llama/use-policy.
 Prohibited Uses
 We want everyone to use Llama 2 safely and responsibly. You agree you will not use, or allow others to use, Llama 2 to:
 1. Violate the law or others’ rights, including to:
 a. Engage in, promote, generate, contribute to, encourage, plan, incite, or further illegal or unlawful activity or content, such as:
 i. Violence or terrorism
 ii. Exploitation or harm to children, including the solicitation, creation, acquisition, or dissemination of child exploitative content or failure to report Child Sexual Abuse Material
 b. Human trafficking, exploitation, and sexual violence
 iii. The illegal distribution of information or materials to minors, including obscene materials, or failure to employ legally required age-gating in connection with such information or materials.
 iv. Sexual solicitation
 vi. Any other criminal activity
 c. Engage in, promote, incite, or facilitate the harassment, abuse, threatening, or bullying of individuals or groups of individuals
 d. Engage in, promote, incite, or facilitate discrimination or other unlawful or harmful conduct in the provision of employment, employment benefits, credit, housing, other economic benefits, or other essential goods and services
 e. Engage in the unauthorized or unlicensed practice of any profession including, but not limited to, financial, legal, medical/health, or related professional practices
 f. Collect, process, disclose, generate, or infer health, demographic, or other sensitive personal or private information about individuals without rights and consents required by applicable laws
 g. Engage in or facilitate any action or generate any content that infringes, misappropriates, or otherwise violates any third-party rights, including the outputs or results of any products or services using the Llama 2 Materials
 h. Create, generate, or facilitate the creation of malicious code, malware, computer viruses or do anything else that could disable, overburden, interfere with or impair the proper working, integrity, operation or appearance of a website or computer system
 2. Engage in, promote, incite, facilitate, or assist in the planning or development of activities that present a risk of death or bodily harm to individuals, including use of Llama 2 related to the following:
 a. Military, warfare, nuclear industries or applications, espionage, use for materials or activities that are subject to the International Traffic Arms Regulations (ITAR) maintained by the United States Department of State
 b. Guns and illegal weapons (including weapon development)
 c. Illegal drugs and regulated/controlled substances
 d. Operation of critical infrastructure, transportation technologies, or heavy machinery
 e. Self-harm or harm to others, including suicide, cutting, and eating disorders
 f. Any content intended to incite or promote violence, abuse, or any infliction of bodily harm to an individual
 3. Intentionally deceive or mislead others, including use of Llama 2 related to the following:
 a. Generating, promoting, or furthering fraud or the creation or promotion of disinformation
 b. Generating, promoting, or furthering defamatory content, including the creation of defamatory statements, images, or other content
 c. Generating, promoting, or further distributing spam
 d. Impersonating another individual without consent, authorization, or legal right
 e. Representing that the use of Llama 2 or outputs are human-generated
 f. Generating or facilitating false online engagement, including fake reviews and other means of fake online engagement
 4. Fail to appropriately disclose to end users any known dangers of your AI system
 Please report any violation of this Policy, software “bug,” or other problems that could lead to a violation of this Policy through one of the following means:
 Reporting issues with the model: github.com/facebookresearch/llama
 Reporting risky content generated by the model: developers.facebook.com/llama_output_feedback
 Reporting bugs and security concerns: facebook.com/whitehat/info
 Reporting violations of the Acceptable Use Policy or unlicensed uses of Llama: LlamaUseReport@meta.com
 """
--- a/Show More
+++ b/Show More